Off-target prediction methods for antigen recognition molecules binding to MHC-peptide targets

JP2025501532A5Pending Publication Date: 2025-12-25REGENERON PHARMACEUTICALS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024536447
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-12-21
Filing Date
2022-12-20
Publication Date
2025-12-25

AI Technical Summary

Technical Problem

Existing engineered antigen recognition molecules, such as TCRs and antibodies, often target off-target cells, leading to severe side effects during clinical trials, necessitating a method to accurately predict and mitigate off-target interactions.

Method used

A computational system and method to predict off-target peptides by analyzing MHC-peptide complexes using computer-based models, determining binding affinities, and identifying amino acid positions involved in interactions with antigen recognition molecules.

Benefits of technology

Enables accurate prediction of off-target peptides, reducing the risk of side effects and optimizing the selection of antigen recognition molecules for targeted therapies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000073_0000
    Figure 00000073_0000
  • Figure 00000073_0001
    Figure 00000073_0001
  • Figure 00000073_0002
    Figure 00000073_0002
Patent Text Reader

Abstract

Described herein are computational systems and methods for predicting the position of amino acids in a target peptide presented in a complex with a major histocompatibility complex (MHC) molecule (MHC-target peptide complex), which are involved in interactions with an antigen recognition molecule that recognizes the MHC-target peptide complex. Described herein are computational systems and methods for estimating the number of off-target peptides of an antigen recognition molecule that recognizes a target peptide presented in a complex with a major histocompatibility complex (MHC) molecule (MHC-target peptide complex). Described herein are computational systems and methods for ranking potential target peptides to reduce off-target toxicity. Such computational systems and methods can streamline the development of effective and well-tolerated antigen recognition molecules for treating disease.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 292,205, filed December 21, 2021, the entire contents of which are incorporated by reference herein.

[0002] Field The present invention relates generally to computational systems and methods for predicting amino acid positions involved in interactions with antigen-recognizing molecules within target peptides presented in complexes with major histocompatibility complex (MHC) molecules (MHC-peptide complexes), and for predicting off-target peptides for antigen-recognizing molecules that recognize target peptides in MHC-peptide complexes. [Background technology]

[0003] background Antigen recognition molecules (such as T cell receptors (TCRs) and antibodies) can identify antigens, including agents recognized by the host's immune system as defined herein. Antigen recognition molecules can assist the immune system in neutralizing antigens by binding to antigenic peptides presented in complexes with major histocompatibility complex (MHC) molecules (MHC-peptide complexes) on the surface of antigen-presenting cells.

[0004] MHC-peptide complexes are presented on the surface of antigen-presenting cells as a result of a cellular process in which MHC genes in antigen cells encode MHC molecules; the MHC molecules then bind antigenic peptides, thereby creating MHC-peptide complexes; and the resulting MHC-peptide complexes are positioned on the cell surface such that a portion of the peptide is presented for binding to antigen-recognizing molecules. Each peptide is composed of a short chain of amino acids, and some amino acids of the peptide in an MHC-peptide complex are bound to the MHC molecule, while at least some of the remaining amino acids are presented and available for binding to antigen-recognizing molecules.

[0005] This cellular process also takes place in cells originating from the body. In general, antigen recognition molecules can distinguish between peptides presented on native cells and peptides presented on antigen-presenting cells, so that normal cells are not attacked by the immune system. Antigen recognition molecules can bind to peptides in the MHC-peptide complex on the surface of antigen-presenting cells to assist the immune system in neutralizing antigens.

[0006] Research has been conducted with the aim of engineering antigen recognition molecules to target cells that would not otherwise be targeted by the above mechanisms. For example, cancer cells are native cells that are not effectively suppressed by the immune system, and research has shown that TCRs, antibodies, and other antigen recognition molecules may be engineered to target cancer-specific MHC-peptide complexes on cancer cells. Targeting treatment with engineered antigen recognition molecules may be effective in neutralizing the intended target cells, but the side effects of treatment may be severe if the engineered antigen recognition molecules attack off-target native cells in addition to the intended target cells. Side effects are often identified during clinical trials, which may result in the death of patients, other adverse effects in patients, and exhaust time and resources for research and development. Therefore, there is a need in the art for methods and systems that can accurately and efficiently predict off-targets for a target peptide of interest, which can help evaluate the risks involved in the target peptide in the target selection process and screen for the most specific antigen recognition molecules. Summary of the Invention

[0007] Abstract One or more computer systems may be configured to perform certain operations or actions by installing software, firmware, hardware, or a combination thereof on the system such that the system operates upon operation. One or more computer programs may be configured to perform certain operations or actions by including instructions that, when executed by a data processing device, cause the device to operate. One general aspect includes a non-transitory computer-readable medium configured to communicate with one or more processors of a computing device. The non-transitory computer readable medium includes instructions for performing the following steps, which can be performed using interleaving steps in various orders: a) receiving as an input a computational representation of a target peptide presented in an MHC-target peptide complex; b) determining the binding affinity of the target peptide to an MHC molecule of the MHC-target peptide complex; c) generating sequences of a plurality of mutated peptides, each associated with a mutation at a respective amino acid position of the target peptide; d) determining the binding affinity of each mutated peptide of the plurality of mutated peptides to the MHC molecule; e) predicting the amino acid positions involved in an interaction with an antigen recognition molecule that recognizes the MHC-target peptide complex based at least in part on a comparison of the binding affinity of each mutated peptide to the binding affinity of the target peptide; and f) providing as an output a representation of the amino acid positions likely to be involved in an interaction with the antigen recognition molecule. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the operations of the method.

[0008] The implementation may include one or more of the following features: The non-transitory computer-readable medium may include instructions that, when executed by a processor, cause the computing device to predict that each of the plurality of positions is involved in an interaction with an antigen recognition molecule that recognizes an MHC-target peptide complex, as determined by a percentile rank value: a) when the percentile rank of said target peptide is 0.5 or less, when the percentile rank of the mutated peptide having a mutation associated with said respective position is less than 1.0, b) when the percentile rank of said target peptide is between 0.5 and 0, when the percentile rank of the mutated peptide having a mutation associated with said respective position is less than 0, or c) when the percentile rank of said target peptide is 0 or more, when the percentile rank of the mutated peptide having a mutation associated with said respective position is less than 4.0. The implementation of the described technology may include hardware, a method or process, or computer software on a computer-accessible medium.

[0009] One general embodiment includes a non-transitory computer-readable medium configured to communicate with one or more processors of a computing device. The non-transitory computer-readable medium includes instructions for performing the following steps, which can be performed in various orders and using interleaving steps: a) receiving as input a computational representation of a target peptide presented in an MHC-target peptide complex; b) predicting all amino acid positions in said target peptide that may be involved in an interaction with an antigen recognition molecule that recognizes said MHC-target peptide complex; d) generating a working list of peptides, wherein within the total pool of predicted or detected peptides of suitable length, the peptides listed in said working list (i) are located at positions corresponding to positions in said target peptide that are involved in an interaction with said antigen recognition molecule. and (ii) generating a working list of peptides such that each peptide comprises at least two amino acids identical to a corresponding amino acid of said target peptide; e) determining the binding affinity of each of the peptides listed in said working list to an MHC molecule of said MHC-target peptide complex; f) filtering said working list to include only peptides having a calculated binding affinity to said MHC molecule that exceeds a first threshold, thereby generating a working list of off-target peptides; and g) providing as output said working list of off-target peptides and / or the number of off-target peptides in said working list. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the operations of the method.

[0010] The implementation may include one or more of the following features. The non-transitory computer-readable medium may include instructions that, when executed by a processor, cause the computing device to estimate the number of peptides in the working list of off-target peptides that are expressed in essential normal tissues; and provide as output the number of off-target peptides expressed in essential normal tissues. The instructions, when executed by a processor, cause the computing device to determine, for each peptide in the working list of peptides, whether such peptide is expressed in essential normal tissues; and filter said working list to include only peptides expressed in essential normal tissues. The instructions, when executed by a processor, cause the computing device to generate a list of potential secondary target peptides that includes peptides that have a calculated binding affinity to MHC molecules higher than a first threshold and have low expression in essential normal tissues. The instructions, when executed by the processor, cause the computing device to calculate a degree of similarity (DoS) score for the peptides in the working list of peptides, said DoS score being based at least in part on the number of amino acids identical to amino acids at corresponding positions of the target peptide, wherein the amino acids of said target peptide are involved in an interaction with the antigen recognition molecule; and filter said working list to include only peptides having a DoS score above a second threshold. Only positions of the target peptide identified as not bound to an MHC molecule are considered in the calculation of the DoS score. The instructions, when executed by the processor, cause the computing device to provide as input a computational representation of an antigen recognition molecule, said antigen recognition molecule capable of binding to an MHC-target peptide complex; determine the binding affinity of said antigen recognition molecule for each likely off-target peptide from the working list and a plurality of MHC-peptide complexes each comprising said MHC molecule; and filter said working list to include only off-target peptides likely to contain a binding motif for said antigen recognition molecule.The instructions, when executed by the processor, cause the computing device to provide as input the off-target peptide expression in essential normal tissues of a particular patient; and provide as output an indication of the off-target effect of said patient. Implementations of the described technology may include hardware, methods or processes, or computer software on a computer-accessible medium.

[0011] One general embodiment includes a non-transitory computer-readable medium configured to communicate with one or more processors of a computing device. The non-transitory computer-readable medium includes instructions for performing the following steps, which can be performed in various orders and using interleaving steps: a) receiving as input a computational representation of a target peptide presented in a complex with a major histocompatibility complex (MHC) molecule (MHC-target peptide complex); b) identifying similar peptides within the total pool of predicted or detected suitable length peptides, (i) located at a position corresponding to a position in said target peptide involved in an interaction with an antigen recognition molecule, and (ii) comprising at least two amino acids identical to the corresponding amino acids of said target peptide; c) determining the binding affinity of each of said identified similar peptides to said MHC molecule; d) identifying off-target peptides based at least in part on the identification of similar peptides having a computed binding affinity to said MHC molecule stronger than a first threshold; and e) providing said off-target peptides as output.

[0012] One general aspect includes a non-transitory computer-readable medium configured to communicate with one or more processors of a computing device. The non-transitory computer-readable medium includes instructions for performing the following steps, which can be performed using interleaving steps in various orders: a) selecting two or more potential target peptides predicted to bind to MHC molecules among disease-associated peptides; b) estimating the number of off-target peptides associated with each of the potential target peptides; and c) ranking the potential target peptides based at least in part on the number of off-target peptides involved with each of the potential target peptides. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the operations of the method.

[0013] The implementation may include one or more of the following features. The non-transitory computer-readable medium may include instructions that, when executed by a processor, cause the computing device to calculate a DoS score for each of the off-target peptides, such that the DoS score indicates the similarity between the respective off-target peptide and the target peptide; and rank the potential target peptides based at least in part on the DoS scores of the off-target peptides associated with each of the potential target peptides. The instructions, when executed by a processor, cause the computing device to calculate a DoS score based at least in part on the number of amino acids of the off-target peptides that are identical to amino acids at corresponding positions of the target peptide, and the amino acids of the target peptides are involved in an interaction with an antigen recognition molecule. Only positions of the target peptide identified as not involved in an interaction with an MHC molecule are considered in the calculation of the DoS score. The instructions, when executed by a processor, cause the computing device to calculate a probability of in vivo toxicity of each potential target peptide based at least in part on the DoS scores of the off-target peptides. The probability of in vivo toxicity of each potential target peptide is based at least in part on the number of highly toxic off-target peptides with DoS scores exceeding a predetermined threshold.The disease-related peptides in step (a) are identified at least in part on the comparison of the expression levels of corresponding mRNA or protein in diseased tissue and essential normal tissue.The implementation of the described technology can include hardware, method or process, or computer software on a computer-accessible medium.

[0014] One general aspect includes a non-transitory computer readable medium comprising a database, said database comprising a plurality of disease associated peptide sequences each associated with an off-target peptide and ranked according to a probability of in vivo toxicity associated with each of said off-target peptides, each of said off-target peptides having approximately equal binding affinity to an MHC molecule as a peptide identified by each of said disease associated peptide sequences. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the operations of the method.

[0015] One general aspect includes a non-transitory computer readable medium comprising a database, said database comprising a plurality of disease associated peptide sequences; a plurality of off-target peptide sequences each associated with a respective disease associated peptide sequence, wherein each disease associated peptide identified by the disease associated peptide sequence has a binding affinity for an MHC molecule that is approximately equal to the binding affinity of the off-target peptide identified by said respective off-target peptide sequence for the MHC molecule. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the operations of the method.

[0016] One general embodiment relates to a method for predicting amino acid positions in a target peptide presented in an MHC-target peptide complex, the amino acid positions being involved in an interaction with an antigen recognition molecule that recognizes the MHC-target peptide complex, the method comprising: a) determining the binding affinity of the target peptide to the MHC molecule; b) generating a plurality of mutated peptides, each associated with a mutation at a respective amino acid position of the target peptide; c) determining the binding affinity of each mutated peptide of the plurality of mutated peptides to the MHC molecule; and d) predicting the amino acid positions involved in an interaction with the antigen recognition molecule based in part on a comparison of the binding affinity of each mutated peptide to the binding affinity of the target peptide.

[0017] In certain embodiments, the binding affinity of a target peptide to an MHC molecule is determined using a computer-based model.

[0018] In certain embodiments, the binding affinity of a target peptide to an MHC molecule is determined by experimental measurements.

[0019] In certain embodiments, the binding affinity of a target peptide to an MHC molecule is measured using the half maximal inhibitory concentration (IC 50 ) value or percentile rank value.

[0020] In certain embodiments, the step of generating a plurality of mutated peptides comprises mutating each amino acid to a glycine amino acid.

[0021] In certain embodiments, the step of generating a plurality of mutated peptides comprises mutating each amino acid to an alanine amino acid.

[0022] In certain embodiments, the method further comprises predicting that each of the plurality of positions is involved in an interaction with an antigen recognition molecule that recognizes an MHC-target peptide complex when the binding affinity of a mutated peptide having a mutation associated with said respective position is approximately equal to the binding affinity of said target peptide.

[0023] In certain embodiments, the method further comprises predicting that each of the plurality of positions, as determined by the percentile rank value, is involved in an interaction with an antigen recognition molecule that recognizes an MHC-target peptide complex when: a) the percentile rank of said target peptide is 0.5 or less, the percentile rank of a mutated peptide having a mutation associated with each of said positions is less than 1.0; b) the percentile rank of said target peptide is between 0.5 and 2.0, the percentile rank of a mutated peptide having a mutation associated with each of said positions is less than 2.0; or c) the percentile rank of said target peptide is greater than 2.0, the percentile rank of a mutated peptide having a mutation associated with each of said positions is less than 4.0.

[0024] In certain embodiments, the method further comprises predicting that each of the plurality of positions is involved in an interaction with an antigen recognition molecule that recognizes an MHC-target peptide complex when both of the following conditions are met: (i) the binding affinity of a mutated peptide having a mutation associated with each of said positions is approximately equal to the binding affinity of said target peptide, and (ii) the amino acid of the target peptide at each of said positions is a non-glycine residue.

[0025] In certain embodiments, the method further comprises verifying that each of the multiple positions in the target peptide is not involved in an interaction with said MHC molecule based on the known structure of the MHC-target peptide complex.

[0026] One general embodiment relates to a method for identifying off-target peptides for an antigen recognition molecule that recognizes a target peptide presented in a complex with a major histocompatibility complex (MHC) molecule (MHC-target peptide complex), comprising the steps of: a) predicting the positions of all amino acids in the target peptide that may be involved in an interaction with the antigen recognition molecule; b) identifying similar peptides of suitable length within the entire pool of predicted or detected peptides, which (i) are located at positions corresponding to positions in the target peptide that are involved in an interaction with the antigen recognition molecule, and (ii) contain at least two amino acids that are identical to the corresponding amino acids of the target peptide; c) determining the binding affinity of each of the identified similar peptides for the MHC molecule; and d) identifying off-target peptides based at least in part on the identification of similar peptides having a calculated binding affinity for the MHC molecule that is stronger than a first threshold value.

[0027] One general embodiment relates to a method for estimating the number of off-target peptides for an antigen recognition molecule that recognizes a target peptide presented in an MHC-target peptide complex, comprising: a) predicting the positions of all amino acids in the target peptide that may be involved in an interaction with the antigen recognition molecule; b) identifying similar peptides of suitable length within the total pool of predicted or detected peptides, (i) located at positions corresponding to positions in the target peptide that are involved in an interaction with the antigen recognition molecule, and (ii) comprising at least two amino acids that are identical to the corresponding amino acids of the target peptide; c) determining the binding affinity of each of the identified similar peptides to the MHC molecule; and d) estimating the number of off-target peptides based at least in part on a count of similar peptides having a calculated binding affinity to the MHC molecule that is stronger than a first threshold value.

[0028] In certain embodiments, step (b) comprises identifying, within the total pool of predicted or detected peptides, similar peptides of suitable length that (i) are located at positions corresponding to positions in said target peptide involved in interaction with said antigen recognition molecule, and (ii) contain at least three amino acids that are identical to the corresponding amino acids of said target peptide.

[0029] In certain embodiments, step (b) comprises identifying, within the total pool of predicted or detected peptides, similar peptides of suitable length that (i) are located at positions corresponding to positions in said target peptide involved in interaction with said antigen recognition molecule, and (ii) contain at least four amino acids that are identical to the corresponding amino acids of said target peptide.

[0030] In certain embodiments, step (b) comprises identifying, within the total pool of predicted or detected peptides, similar peptides of suitable length that (i) are located at positions corresponding to positions in said target peptide involved in interaction with said antigen recognition molecule, and (ii) contain at least 5 amino acids that are identical to the corresponding amino acids of said target peptide.

[0031] In certain embodiments, the binding affinity in step (c) is determined by determining a percentile rank value, wherein the number of said off-target peptides is estimated in step (d) based at least in part on the count of similar peptides having a percentile rank value of less than 2.0.

[0032] In certain embodiments, the method further comprises including among the off-target peptides only peptides that are expressed in essential normal tissues.

[0033] In certain embodiments, the method further comprises, for each identified off-target peptide, determining whether such peptide is expressed in an essential normal tissue.

[0034] In certain embodiments, peptide expression is determined based on the expression levels of the corresponding mRNA or protein.

[0035] In certain embodiments, peptide expression is determined based on mass spectrometry data.

[0036] In certain embodiments, the method further comprises including among the off-target peptides only similar peptides that have a calculated binding affinity to the MHC molecule stronger than a first threshold and detectable expression in requisite normal tissues.

[0037] In certain embodiments, the method further comprises calculating a degree of similarity (DoS) score for the similar peptides, said DoS score being based at least in part on the number of amino acids of each similar peptide that are identical to amino acids at corresponding positions of the target peptide, said amino acids of the target peptide being involved in an interaction with the antigen recognition molecule.

[0038] In certain embodiments, only those positions of the target peptide that are identified as not bound to the MHC molecule are considered in calculating the DoS score.

[0039] In certain embodiments, the method further comprises including among the off-target peptides only similar peptides having a DoS score higher than a second threshold.

[0040] In certain embodiments, the step of predicting the positions of amino acids in the target peptide involved in interaction with the antigen recognition molecule in step (a) is carried out using a method for estimating the number of off-target peptides for an antigen recognition molecule that recognizes a target peptide presented in an MHC-target peptide complex as disclosed above.

[0041] In certain embodiments, the entire pool of detected peptides in step (b) is based on mass spectrometry data.

[0042] In certain embodiments, the MHC molecule is a class I MHC molecule and the predicted or detected peptides in step (b) are between 8 and 12 amino acids in length.

[0043] In certain embodiments, the MHC molecule is a class I MHC molecule and the target peptide is 8-12 amino acids in length.

[0044] In certain embodiments, the antigen recognition molecule is a T cell receptor (TCR), a chimeric antigen receptor (CAR), an antibody, or an antigen-binding fragment thereof.

[0045] One general embodiment relates to a method of ranking potential target peptides to reduce off-target toxicity, the method comprising: a) selecting two or more potential target peptides among disease-associated peptides predicted to bind to an MHC molecule; b) estimating the number of off-target peptides associated with each of the potential target peptides; and c) ranking the potential target peptides based at least in part on the number of off-target peptides associated with each of the potential target peptides.

[0046] In certain embodiments, the number of off-target peptides in step (b) is estimated using the methods described herein.

[0047] In certain embodiments, the method further comprises a step of ranking said potential target peptides such that potential target peptides with fewer associated off-target peptides are selected for further analysis and / or used for generation and / or testing of antigen-recognizing molecules.

[0048] In certain embodiments, the method further comprises calculating a degree of similarity (DoS) score for each of the off-target peptides, such that the DoS score indicates the similarity between each off-target peptide and the target peptide; and ranking said potential target peptides based at least in part on the DoS scores of the off-target peptides associated with each of the potential target peptides.

[0049] In certain embodiments, the calculation of the DoS score is based at least in part on the number of amino acids of the off-target peptide that are identical to amino acids at corresponding positions of the target peptide, said amino acids of the target peptide being involved in an interaction with an antigen recognition molecule.

[0050] In certain embodiments, only positions of the target peptide identified as not involved in interactions with MHC molecules are considered in calculating the DoS score.

[0051] In certain embodiments, the method further comprises calculating a probability of in vivo toxicity for each potential target peptide based at least in part on the DoS scores of the off-target peptides.

[0052] In certain embodiments, the probability of in vivo toxicity for each potential target peptide is based at least in part on the number of highly toxic off-target peptides having a DoS score above a predetermined threshold.

[0053] In certain embodiments, the method further comprises the step of ranking the potential target peptides such that those with lower toxicity are prioritized.

[0054] In certain embodiments, the disease-associated peptides in step (a) are identified based at least in part on a comparison of expression levels of corresponding mRNA or protein in diseased and requisite normal tissues.

[0055] In certain embodiments, the method further comprises providing a ranking of the potential target peptides in a database.

[0056] In certain embodiments, the method further comprises providing a list of off-target peptides associated with each of the potential target peptides in the database.

[0057] In another aspect, provided herein is an in vitro method of assessing off-target effects of an antigen recognition molecule, comprising: a) contacting said antigen recognition molecule with a target peptide presented in a complex with a major histocompatibility complex (MHC) molecule (MHC-target peptide complex); b) contacting said antigen recognition molecule with one or more off-target peptides related to said target peptide, wherein each of said off-target peptides is presented in a complex with the same MHC molecule as in (a) (MHC-off-target peptide complex); and c) determining and comparing the binding level of the antigen recognition molecule to the MHC-target peptide complex and each of said MHC-off-target peptide complexes.

[0058] In another aspect, provided herein is an in vitro method of assessing an off-target effect of an antigen recognition molecule, comprising: a) contacting said antigen recognition molecule with one or more off-target peptides related to a target peptide recognized by said antigen recognition molecule, wherein each of said off-target peptides is presented in a complex with a major histocompatibility complex (MHC) molecule (MHC-off-target peptide complex); and b) determining the binding level of said antigen recognition molecule to each of said MHC-off-target peptide complexes.

[0059] In some embodiments of the above in vitro methods, the method may further comprise determining that an antigen recognition molecule is likely to have an off-target effect if the antigen recognition molecule detectably binds to at least one MHC-off-target peptide complex, wherein the off-target peptide is expressed in an essential normal tissue.

[0060] In another aspect, provided herein is a method of selecting an antigen recognition molecule, the method comprising: a) contacting a plurality of antigen recognition molecules with a target peptide presented in a complex with a major histocompatibility complex (MHC) molecule (MHC-target peptide complex); b) contacting the same plurality of antigen recognition molecules with one or more off-target peptides associated with the target peptide, wherein each of the off-target peptides is presented in a complex with the same MHC molecule as in (a) (MHC-off-target peptide complex); c) selecting one or more antigen recognition molecules based at least in part on the number of MHC-off-target peptide complexes detectably bound by each of the antigen recognition molecules; and d) optionally repeating steps (a)-(c) using the one or more selected antigen recognition molecules.

[0061] In some embodiments, one or more selected antigen recognition molecules detectably bind to no more than 5 (e.g., no more than 4, no more than 3, no more than 2, or no more than 1) MHC-off-target peptide complexes, where said off-target peptides are expressed in requisite normal tissue.

[0062] In some embodiments, the one or more selected antigen recognition molecules do not detectably bind to any MHC-off-target peptide complexes, where said off-target peptides are expressed in requisite normal tissues.

[0063] In some embodiments, the plurality of antigen recognition molecules is present in a library, hi some embodiments, the library is a phage display library or a yeast library.

[0064] In various embodiments, the MHC-peptide complex is immobilized on a solid support. In various embodiments, the MHC-peptide complex is present on an antigen presenting cell.

[0065] In various embodiments, the level of binding is determined by detecting the amount of binding of an antigen recognition molecule to the MHC-peptide complex.

[0066] In various embodiments, the methods are performed in a high-throughput format (eg, 96-well plates).

[0067] In another aspect, provided herein is a method of enriching a sample with antigen recognition molecules that specifically bind to a target peptide, the method comprising: a) contacting a sample comprising a plurality of antigen recognition molecules with the target peptide in the presence of one or more off-target peptides associated with the target peptide, wherein each of the target peptide and the one or more off-target peptides are presented in a complex with a major histocompatibility complex (MHC) molecule (MHC-target peptide complex or MHC-off-target peptide complex); b) enriching the sample by isolating the antigen recognition molecules bound to the MHC-target peptide complex; and c) optionally repeating steps (a)-(b) using the enriched sample.

[0068] In some embodiments of the above method, the MHC-target peptide complex is present on an antigen-presenting cell, and the MHC-off-target peptide complex is not present on the antigen-presenting cell. In some embodiments, the MHC-target peptide complex is immobilized on a solid support, and the MHC-off-target peptide complex is not immobilized on a solid support. In some embodiments, the MHC-target peptide complex is labeled, and the MHC-off-target peptide complex is not labeled or differentially labeled as the MHC-target peptide complex.

[0069] In various embodiments of the above methods, one or more off-target peptides are identified using the methods described herein.

[0070] In various embodiments of the above method, the antigen recognition molecule is a T cell receptor (TCR), a chimeric antigen receptor (CAR), an antibody, or an antigen-binding fragment thereof.In some embodiments, the antigen recognition molecule is in solution.In some embodiments, the antigen recognition molecule is on a cell.

[0071] In yet a further aspect, provided herein is a library comprising at least two off-target peptides identified using the methods described herein. In some embodiments, the library of the present disclosure comprises one or more peptides each selected from the amino acid sequences of SEQ ID NOs: 1-7, 9, and 11-74, or any combination thereof. In some embodiments, the at least two off-target peptides are each present in a complex with a major histocompatibility complex (MHC) molecule.

[0072] BRIEF DESCRIPTION OF THE DRAWINGS The above and further aspects of the present invention should be understood with reference to the drawings, in which similar elements in different drawings are numbered the same. The drawings, which are not necessarily to scale, depict selected embodiments and are not intended to limit the scope of the invention. The detailed description illustrates the principles of the invention, but does not limit it. The description describes several embodiments, adaptations, modifications, alternatives, and uses of the invention, including what is presently believed to be the best mode for carrying out the invention, clear enough to enable a person skilled in the art to make and use the invention. [Brief description of the drawings]

[0073] [Figure 1] FIG. 1A shows a flow diagram of an exemplary embodiment of the peptide in groove similarity prediction (PIGSPRED) method.

[0074] FIG. 1B shows a flow diagram of an exemplary embodiment of a method for predicting the position of an amino acid in a target peptide presented in a complex with a major histocompatibility complex (MHC) molecule (MHC-target peptide complex), which can be performed as part of the PIGSPRED method.

[0075] [Diagram 2] FIG. 2 shows a flow diagram of an embodiment of a method for calculating expected off-target toxicity associated with an MHC-target peptide complex.

[0076] [Diagram 3] FIG. 3 shows a flow diagram of an embodiment of a method for prioritizing potential target peptides to reduce off-target toxicity.

[0077] [Figure 4] FIG. 4 shows a block diagram of an exemplary embodiment of a computational system configured to output a list of off-target peptides and / or metrics of off-target toxicity, said system including an exemplary embodiment of the PIGSPRED engine.

[0078] [Diagram 5] FIG. 5 shows a block diagram of an exemplary embodiment of a computational system configured to output a list of low-risk peptide targets, a list of off-target peptides for potential targets, and / or metrics of off-target toxicity for potential targets, the system including an exemplary embodiment of a target ranking engine.

[0079] [Figure 6] FIG. 6 shows a block diagram of an exemplary embodiment of a target toxicity database.

[0080] [Figure 7] FIG. 7 illustrates a block diagram of an embodiment of a computing device.

[0081] [Figure 8] FIG. 8 shows a block diagram of an embodiment of a computation network.

[0082] [Figure 9] FIG. 9 illustrates cellular functions for the exemplary embodiments presented herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0083] definition Unless otherwise defined, all technical and scientific terms used herein have the meaning commonly understood by one of ordinary skill in the relevant art.

[0084] The singular forms "a," "an," and "the" include the plural unless the context clearly dictates otherwise. Thus, for example, reference to a "method" includes one or more methods, and / or steps of the type described herein and / or that will be apparent to those skilled in the art upon reading this disclosure.

[0085] The term "about" or "approximately" includes within a range of values ​​having meaning. The allowable variation encompassed by the term "about" or "approximately" depends on the particular system under study, and this variation can be readily discerned by one of ordinary skill in the art.

[0086] The terms "major histocompatibility complex" and "MHC" refer to the terms "human leukocyte antigen" or "HLA" (of which the latter two are commonly used terms for human MHC molecules), naturally occurring MHC molecules (e.g., MHC class I molecules, which include MHC class I α (heavy) chains and β2 microglobulin; MHC class II molecules, which include MHC class II α chains and MHC class II β chains), the individual chains of MHC molecules (e.g., MHC class I α (heavy) chains, MHC class II α chains, and MHC class II β chains), the chains of such MHC molecules, and the MHC class II molecules themselves (e.g., MHC class I α (heavy) chains, MHC class II α chains, and MHC class II β chains). MHC class I molecules include individual subunits of MHC class I (e.g., α1, α2, and / or α3 subunits of the MHC class I α chain, α1-α2 subunits of the MHC class II α chain, β1-β2 subunits of the MHC class II β chain), as well as portions (e.g., peptide-binding portions, e.g., peptide-receiving grooves), variants, and various derivatives (including fusion proteins) thereof, where such portions, variants, and derivatives retain the ability to display antigenic peptides for recognition by a T cell receptor (TCR), e.g., an antigen-specific TCR. MHC class I molecules contain a peptide-receiving groove formed by the α1 and α2 domains of the heavy chain that can accommodate peptides of approximately 8-10 amino acids. Despite the fact that both MHC classes bind to a core of approximately 9 amino acids (e.g., 5-17 amino acids) within a peptide, a wider range of peptide lengths is tolerated because there is no constraint on the MHC class II peptide-receiving groove (the α1 domain of a class II MHC α polypeptide associated with the β1 domain of a class II MHC β polypeptide). Peptides that bind to MHC class II usually vary between 13 and 17 amino acids in length, but shorter or longer lengths are not uncommon. As a result, peptides move within the MHC class II peptide-accommodating groove, and the 9-mer that resides directly within the groove changes. In some embodiments, the MHC-peptide complexes described herein can be MHC-peptide complexes derived from non-human animals. In other embodiments, the MHC-peptide complexes described herein can include HLA-peptide complexes, i.e., MHC-peptide complexes derived from humans.

[0087] The term "non-human animal" refers to any vertebrate that is not a human. In some embodiments, the non-human animal is a cyclostome, a bony fish, a cartilaginous fish (e.g., a shark or a ray), an amphibian, a reptile, a mammal, and a bird. In some embodiments, the non-human animal is a mammal. In some embodiments, the non-human mammal is a primate, a goat, a sheep, a pig, a dog, a cow, or a rodent. In some embodiments, the non-human animal is a rodent, such as a rat or a mouse.

[0088] The term "antigen" refers to any agent (e.g., a protein, peptide, polysaccharide, glycoprotein, glycolipid, nucleotide, portion thereof, or combination thereof) that is recognized by the host's immune system when introduced into an immunocompetent host and elicits an immune response by the host. T cell receptors recognize peptides presented in association with the major histocompatibility complex (MHC) as part of the immune synapse. The peptide-MHC (pMHC) complex is recognized by the TCR, with the peptide (antigenic determinant) and the TCR idiotype providing the specificity of the interaction. Thus, the term "antigen" encompasses peptides presented in association with MHC (e.g., peptide-MHC complexes). Peptides displayed on MHC may also be referred to as "epitopes" or "antigenic determinants." The terms "peptide," "antigenic determinant," "epitope," and the like, include those naturally presented by antigen-presenting cells (APCs), but can also be any desired peptide, so long as it is recognized by the immune cells of an animal when appropriately presented to the cells of the immune system, for example. For example, a peptide having an artificially prepared amino acid sequence can also be used as an epitope.

[0089] The term "antigen recognition molecule" refers to any molecule that can recognize an antigen as defined above. Antigen recognition molecules can include, but are not limited to, T cell receptors (TCRs), antibodies, antibody fragments, or chimeric antigen receptors (CARs).

[0090] "MHC-peptide complex", "peptide-MHC complex", "pMHC complex", and "peptide-in-groove" include the following:

[0091] (i) an MHC molecule, e.g., a human and / or non-human animal MHC molecule, or a portion thereof (e.g., its peptide-accommodating groove, e.g., its extracellular portion), and

[0092] (ii) An antigenic peptide, in which the MHC molecule and the antigenic peptide are complexed in such a manner that the pMHC complex is capable of specific binding to a T cell receptor. pMHC complexes include pMHC complexes expressed on the cell surface and soluble pMHC complexes.

[0093] "HLA-peptide complex," "peptide-HLA complex," and "pHLA complex," and the like, refer to an MHC-peptide complex in which the MHC molecule is a human leukocyte antigen (HLA) molecule.

[0094] The terms "antibody", "antibodies", "immunoglobulins", and "binding proteins", and the like, refer to monoclonal antibodies, multispecific antibodies, human antibodies, humanized antibodies, chimeric antibodies, single chain Fvs (scFvs), single chain antibodies, Fab fragments, F(ab') fragments, disulfide-linked Fvs (sdFvs), intrabodies, minibodies, diabodies, and anti-idiotypic (anti-Id) antibodies (including, for example, anti-Id antibodies against antigen-specific TCRs), and epitope-binding fragments of any of the above. The terms "antibody" and "antibodies" also refer to covalent diabodies (such as those disclosed in U.S. Patent Publication No. 20070004909, which is incorporated by reference in its entirety) and Ig-DARTS (such as those disclosed in U.S. Patent Publication No. 20090060910, which is incorporated by reference in its entirety). A pMHC binding protein refers to an antigen binding protein, immunoglobulin, or antibody, etc., that specifically binds to a pMHC complex.

[0095] "Individual" or "subject" or "animal" refers to humans, domestic animals (e.g., cats, dogs, cows, horses, sheep, pigs, etc.) and experimental animal models of disease (e.g., mice, rats). In one embodiment, the subject is a human.

[0096] As used herein, the term "protein" encompasses all kinds of naturally occurring and synthetic proteins (including protein fragments of all lengths), fusion proteins and modified proteins (including, but not limited to, glycoproteins), as well as all other types of modified proteins (e.g., proteins obtained by phosphorylation, acetylation, myristoylation, palmitoylation, glycosylation, oxidation, formylation, amidation, polyglutamylation, ADP-ribosylation, pegylation, biotinylation, etc.).

[0097] The terms "nucleic acid" and "nucleotide" encompass both DNA and RNA, unless otherwise specified.

[0098] The term "library" refers to an isolated population of at least two elements that differ from each other in at least one state. For example, a "peptide library" is a population of at least two peptides that may differ from each other in at least one amino acid. As another example, a "pMHC complex library" is a population of pMHC complexes that may differ from each other in at least one amino acid in the peptide or at least one MHC polypeptide. The elements of the library are isolated from elements of a similar type that are not part of the library (e.g., peptides of a peptide library are isolated from peptides that are not part of the library). The library may exist in vitro or ex vivo.

[0099] The term "administration" and the like refers to and includes administration of a composition (e.g., an antigen-recognizing molecule) to a subject or system (e.g., a cell, an organ, a tissue, an organism, or a related component or set of components thereof). Those skilled in the art will recognize that the route of administration may vary depending, for example, on the subject or system to which the composition is administered, the nature of the composition, the purpose of administration, and the like. For example, in certain embodiments, administration to an animal subject (e.g., to a human or rodent) may be bronchial (including bronchial instillation), oral, intestinal, interdermal, intraarterial, intradermal, intragastric, intramedullary, intramuscular, intranasal, intraperitoneal, intrathecal, intravenous, intraventricular, mucosal, nasal, oral, rectal, subcutaneous, sublingual, topical, tracheal (including intratracheal instillation), transdermal, vaginal, and / or vitreous. In some embodiments, administration may require intermittent administration. In some embodiments, administration may require continuous administration (e.g., perfusion) for at least a selected period of time.

[0100] The term "essential normal tissue" refers to the tissue of a patient in which the activity of a given antigen-recognizing molecule administered for the treatment of a disease may result in unacceptable side effects. The list of tissues considered essential and normal will vary depending on the disease being treated and the risk associated with the disease itself (e.g., the list will be smaller for life-threatening diseases than for non-life-threatening diseases). For example, but not limited to, when treating a life-threatening cancer, tissue types that may be considered non-essential may include breast, ovary, and testis. The list of tissues considered essential and normal will also vary depending on the likelihood that a given antigen-recognizing molecule will reach such tissue. For example, in the case of an antigen-recognizing molecule that does not penetrate the blood-brain barrier of a patient being treated for a disease, the brain may not be included in the list of essential normal tissues.

[0101] The terms "component," "engine," "module," "system," "server," "processor," and "memory," etc., are intended to include one or more computer-related units (such as, but not limited to, hardware, firmware, a combination of hardware and software, software, or software in execution). For example, a component may be, but not limited to, a process, an object, an executable, a thread of execution, a program, and / or a computer running on a processor. By way of illustration, both an application running on a computing device and a computing device may be a component. One or more components may reside within a process and / or thread of execution, and a component may be localized on one computer and / or distributed among two or more computers. Furthermore, these components may execute from various computer-readable media having various data structures stored thereon. A component may communicate via local and / or remote processes, for example, according to signals having one or more data packets (such as data exchanged from one component with another component in a local system, in a distributed system, and / or over a network such as the Internet with other systems via signals).

[0102] The term "connected" means that one feature, characteristic, structure, or property is directly connected or in communication with another feature, characteristic, structure, or property.

[0103] The term "coupled" means that one feature, characteristic, structure, or property is directly or indirectly connected or in communication with another feature, characteristic, structure, or property.

[0104] The terms "comprising" or "containing" or "including" mean that at least the named elements or method steps are present in an item or method, but do not exclude the presence of other elements or method steps, even if such other elements or method steps do not have the same function as the one named.

[0105] As used herein, unless otherwise specified, the use of the ordinal adjectives "first," "second," "third," etc. to describe a general object indicates only that different instances of a similar object are being referred to, and is not intended to imply that the described objects must be in a given order, either temporally, spatially, historically, or in any other manner.

[0106] In this description, numerous specific details are described. However, it should be understood that implementations of the disclosed technology may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been described in detail so as not to obscure an understanding of this description. References to "one embodiment," "an embodiment," "several embodiments," "exemplary embodiments," "various embodiments," "one implementation," "an implementation," "exemplary implementation," "various implementations," "several implementations," and the like indicate that implementations of the disclosed technology so described may include a particular feature, structure, or characteristic, but not every implementation necessarily includes the particular feature, structure, or characteristic. Furthermore, repeated use of the phrase "in one implementation" may, but does not necessarily, refer to the same implementation.

[0107] Detailed Description Some embodiments presented herein relate to the identification of off-target peptides that are similar to the intended target peptide of an MHC-target peptide complex, such that an antigen recognition molecule engineered for the intended target MHC-target peptide complex is likely to also target the off-target peptide. In some examples, the identification of off-target peptides can include predicting the amino acid positions within the target peptide presented in the MHC-target peptide complex that are available to participate in an interaction with the antigen recognition molecule. In some embodiments, the identified off-target peptides can be used to predict the probability of in vivo toxicity of a treatment targeting the target peptide. In some embodiments, off-target peptides can be identified for several potential target peptides, and the target peptides can be ranked at least in part based on their associated off-target peptides. In some embodiments, the identification of off-target peptides can be independent of the antigen recognition molecule, and thus the identification of off-target peptides, prediction of in vivo toxicity, and / or ranking of potential target peptides can be used to guide the development of engineered antigen recognition molecules to target target peptides with fewer identified off-target peptides, lower predicted probability of in vivo toxicity, and / or favorable ranking. Some embodiments disclosed herein include a computing system, engine, module, device, and / or network configured to perform most of the steps related to the above-mentioned embodiments. The output of such a computing system can be used to provide information for the development of antigen recognition molecules, screening of patients undergoing clinical trials, treatment of individual patients, and other applications understood by those skilled in the relevant art following the teachings herein. One objective of some embodiments shown herein is to avoid side effects that would otherwise be identified during clinical trials, thereby reducing the number of patient deaths, reducing other adverse effects of patients, and reducing the consumption of time and resources in research and development.

[0108] MHC molecules are generally classified into two categories: class I and class II MHC molecules. MHC class I molecules, also referred to herein as α chains, are integral membrane proteins that contain a glycoprotein heavy chain with three extracellular domains (i.e., α1, α2, and α3) and two intracellular domains (i.e., transmembrane domain (TM) and cytoplasmic domain (CYT)). The heavy chain is non-covalently associated with a soluble subunit called β2 microglobulin (β2m or β2M). MHC class II molecules or proteins are heterodimeric integral membrane proteins that contain one α chain and one β chain non-covalently associated. The α chain has two extracellular domains (α1 and α2), and two intracellular domains (TM and CYT domains). The β chain contains two extracellular domains (β1 and β2), and two intracellular domains (TM and CYT domains).

[0109] The domain organization of class I and class II MHC molecules results in the formation of an antigenic determinant binding site (e.g., the peptide-binding portion or peptide-binding groove of the MHC molecule). The peptide-binding groove refers to the portion of the MHC protein that forms a cavity to which a peptide (e.g., an antigenic determinant) can bind. The conformation of the peptide-binding groove can change upon binding to a peptide to allow proper alignment of amino acid residues important for binding of the TCR to the peptide-MHC (pMHC) complex.

[0110] In some embodiments, the MHC molecule comprises a fragment of an MHC chain sufficient to form a peptide-accommodating groove. For example, the peptide-accommodating groove of a class I protein can comprise a portion of the α1 and α2 domains of the heavy chain capable of forming two β-pleated sheets and two α-helices. The inclusion of a portion of the β2 microglobulin chain stabilizes the MHC class I molecule. While in most versions of MHC class II molecules, the interaction of the α and β chains can occur in the absence of peptide, the MHC class II two-chain molecule is unstable until the accommodation groove is filled with peptide. The peptide-accommodating groove of a class II protein can comprise a portion of the α1 and β1 domains capable of forming two β-pleated sheets and two α-helices. A first portion of the α1 domain forms a first β-pleated sheet and a second portion of the α1 domain forms a first helix. A first portion of the β1 domain forms a second β-pleated sheet and a second portion of the β1 domain forms a second helix. X-ray crystallographic structures of class II proteins with peptides engaged in the accommodation groove of the protein show that one or both ends of the engaged peptide can protrude beyond the MHC protein (Brown et al., pp. 33-39, 1993, Nature, Vol. 364; incorporated herein by reference in its entirety). Thus, the ends of the α1 and β1 α-helices of class II form an open cavity such that the ends of the peptides bound to the accommodation groove are not buried in the cavity. Furthermore, X-ray crystallographic structures of class II proteins show that the N-terminus of the MHC β-chain apparently protrudes from the side of the MHC protein in an unstructured manner, since the first four amino acid residues of the β-chain cannot be assigned by X-ray crystallography. Numerous human and other mammalian MHCs are well known in the art.

[0111] In some embodiments, the MHC molecule may be a human HLA molecule selected from the group consisting of HLA-A, HLA-B, HLA-C, HLA-E, HLA-F, and HLA-G. A list of commonly used HLA alleles is provided in Shankarkumar et al. ((2004) The Human Leukocyte Antigen (HLA) System, Int. J. Hum. Genet. 4(2):91-103), which is incorporated by reference in its entirety. Shankarkumar et al. also provide a brief description of the HLA nomenclature used in the art. Further information regarding HLA nomenclature and various HLA alleles can be found in Holdsworth et al. (2009) The HLA dictionary 2008: a summary of HLA-A, -B, -C, -DRB1 / 3 / 4 / 5, and DQB1 alleles and their association with serologically defined HLA-A, -B, -C, -DR, and -DQ antigens, Tissue Antigens 73:95-170 and a recent update by Marsh et al. (2010) Nomenclature for factors of the HLA system, 2010, Tissue Antigens 75:291-455 (each of these publications is incorporated herein by reference in its entirety). In some embodiments, the MHC I or MHC II polypeptide may be derived from any functional human HLA-A, B, C, DR, or DQ molecule. In one embodiment, the HLA molecule is HLA-A2 (HLA-A * In another embodiment, the HLA molecule is encoded by HLA-A1 (HLA-A2:01 allele, etc.). * 01:01 allele, etc.

[0112] Targeting peptide-MHC (pMHC) complexes specifically expressed on cells, such as cancer cells, via antibody-based or cell-based therapeutic approaches can be an effective method of destruction of such cells. However, potential off-targets associated with these pMHC complexes can often lead to off-target toxicity. The present disclosure provides, among other things, a method named PIGSPRED (peptide in groove similarity prediction) that is useful for predicting such off-targets.

[0113] FIG. 1A shows a flow diagram of an exemplary embodiment of the Peptides in Groove Similarity Prediction (PIGSPRED) method 100.

[0114] In step 102, an MHC-target peptide complex is provided as an input to the PIGSPRED method. The MHC-target peptide complex comprises a target peptide and an MHC molecule.

[0115] In step 104, peptide positions (amino acids) important for antigen recognition molecule binding are identified. The target peptide contains some amino acids that are bound to the MHC molecule and some amino acids that are available for binding to the antigen recognition molecule. When evaluating the similarity / homology of a peptide to a target peptide, the similarity can only be evaluated for positions of the peptide that are available for binding to the antigen recognition molecule (i.e., positions of amino acids within the peptide). These available positions can be confirmed by analyzing experimental structures of MHC-target peptide complexes with specific antigen recognition molecules, where such experimental structures are typically derived using crystallography or cryoEM techniques. Thus, peptide positions important for antigen recognition molecule binding can be identified by analyzing experimental structures; however, the specific antigen recognition molecule must be known, and such experimental structures can be difficult to obtain. When developing targeted therapy, at the initial target selection stage, when potential risks associated with the target are evaluated, specific antigen recognition molecules for the target are not available. The Immune Epitope Database (IEDB) is a freely available resource that catalogs experimental data on antibody and T cell epitopes studied in humans, non-human primates, and other animal species in the context of infectious diseases, allergy, autoimmunity, and transplantation. The IEDB also hosts tools that aid in epitope prediction and analysis. Similar databases and methodologies for characterizing epitopes in such databases can be used to obtain experimental structures of MHC-peptide complexes.

[0116] FIG. 1B illustrates steps of the PIGSPRED method 100 that may be performed in step 104. The steps illustrated in FIG. 1B illustrate an exemplary embodiment of a method 104 for predicting the position of an amino acid in a target peptide presented in an MHC-target peptide complex. Method 104 may be performed without relying on the structure of the antigen recognition molecule, thus providing an alternative to the analysis of the experimental structure of an MHC-target peptide complex with a specific antigen recognition molecule when the specific antigen recognition molecule is unknown. The positions identified by the computer in step 104 may be verified by experimental data such as mass spectrometry data or otherwise available in a database similar to the IEDB.

[0117] At step 110, the binding affinity of the MHC molecule to the target peptide is predicted. In some embodiments, the binding affinity can be calculated using a computer-based model. As a non-limiting example, a commercially available tool called NetMHCpan, which uses a machine learning model, utilizes a computer-based binding affinity model that may be suitable for computer calculation of binding affinity. Any version of NetMHCpan may be used (e.g., NetMHCpan4.0 or NetMHCpan4.1). Other non-limiting examples include MHCflurry (see, e.g., O'Donnell et al., Cell Syst. 2018 Jul 25; 7(1): 129-132.e4, which is incorporated herein by reference in its entirety) and MixMHCPred (see, e.g., Boehm et al., BMC Bioinformatics volume 20, Article number: 7 (2019), which is incorporated herein by reference in its entirety).

[0118] Predicting binding affinity estimates how tightly a peptide will bind to a specific MHC molecule, expressed as a predicted IC in the nanomolar (nM) range. 50 The predicted IC is measured either as a percentile or as a percentile rank prediction. 50The experimental IC value measures the concentration of a competing peptide required to displace 50% of the binding of the peptide to the MHC molecule. 50 In some embodiments, a peptide is considered a binder if its predicted binding affinity is 500 nM or less; if it is 50 nM or less, it is considered a strong binder. The percentile rank value (e.g., %Rank_BA in NetMHCpan) represents the IC20 between different MHC molecules. 50 Normalize the values. The ranks are then normalized to a series of (approximately 10 5 (1) Predicted IC of naturally occurring random peptides 50 Predicted IC of peptides against values 50 The percentile rank is calculated by comparing the percentile ranks (see, e.g., Jurtz V et al., J Immunol. 2017, incorporated herein by reference in its entirety). In some embodiments, a peptide is considered a binder if its predicted percentile rank is 2 or less; if it is 0.5 or less, it is considered a strong binder.

[0119] Computational mutagenesis can be performed in steps 112 and 114, where in step 112 the target peptide is mutated at a single amino acid position, resulting in several mutated peptides, each with a single amino acid mutation from the target peptide. The mutated peptides are each associated at their respective peptide positions (i.e., positions of mutation). Preferably, the target peptide is mutated at every amino acid position, resulting in a number of mutated peptides equal to the number of amino acids in the target peptide. In some embodiments, the mutated peptides can include a single amino acid mutation to a glycine or alanine amino acid. In some embodiments, the target peptide is mutated at a single amino acid position, where said amino acid is replaced with a glycine, and positions in the target peptide that have glycine are skipped, resulting in several mutated peptides corresponding to the mutation of the target peptide at non-glycine positions.

[0120] In step 114, the binding affinity of each mutated peptide to the MHC molecule is predicted. In some embodiments, the binding affinity can be calculated using a computer-based model that is the same as or similar to the computer-based model utilized in step 110.

[0121] In step 116, the binding affinity of the target peptide to the MHC molecule is compared to the binding affinity of each of the mutated peptides to the MHC molecule. Peptide positions associated with the mutated peptides that do not lose binding affinity to the MHC molecule (compared to the binding affinity of the target peptide to the MHC molecule) are flagged. These flagged peptide positions are identified as not involved in binding to the MHC molecule and are therefore free to interact with antigen recognition molecules.

[0122] In some embodiments, when the binding affinity of the mutated peptide with the mutations related to each of the aforementioned positions is approximately equal to the binding affinity of the target peptide to the MHC molecule, each position related to the mutated peptide is predicted to be involved in the interaction with the antigen recognition molecule that recognizes the MHC-target peptide complex, or is flagged as a free position. Percentile rank values ​​from NetMHCpan (e.g., NetMHCpan4.0 or NetMHCpan4.1) may be used. If the percentile rank of the mutated peptide is below a certain threshold, the mutated position is determined to be a position that has lost binding affinity to the MHC molecule, where the lower the percentile rank, the higher the binding affinity. Using the predicted percentile rank to quantify the binding affinity, in some embodiments, if the rank of the target peptide is 0.5 or less, the threshold is 1.0. If the rank of the target peptide is greater than 0.5 and less than or equal to 2.0, the threshold is 2.0. If the rank of the target peptide is greater than 2.0, the threshold is 4.0. In other embodiments, the difference in binding affinity (e.g., percentile rank) between the mutated peptide and the target peptide may be used to identify positions involved in binding to MHC molecules or free to interact with antigen recognition molecules. Positions where the mutation results in a large loss of binding affinity (e.g., loss greater than a threshold) may be determined to be involved in binding to MHC molecules, and positions where the mutation does not result in a large loss of binding affinity may be determined to be free to interact with antigen recognition molecules. For example, if the target peptide rank is 0.5 or less, the difference in rank (loss) threshold may be 1.0. If the target peptide rank is greater than 0.5 and less than or equal to 2.0, the difference in rank threshold may be 1.5. If the target peptide rank is greater than 2.0, the difference in rank threshold may be 2.0.

[0123] Among these flagged free positions, positions containing non-glycine amino acids are identified as important positions for antigen recognition molecule binding in step 118. In some embodiments, a position is considered important for antigen recognition molecule binding if the loss of MHC molecule binding is not significant when the position is mutated and the amino acid at the position is a non-glycine amino acid. In embodiments in which the target peptide is mutated at a single amino acid position by substituting the amino acid with glycine in step 112 and the positions of the target peptide having glycine are skipped, step 118 can be omitted since all of the mutated peptides resulting from step 112 correspond to positions containing non-glycine amino acids by definition.

[0124] In some embodiments, when the structure of the MHC-target peptide complex bound to the antigen recognition molecule is known, the binding motif in the peptide sequence involved in the interaction with the antigen recognition molecule can be used as a comparison to verify that the positions important for antigen recognition molecule binding are correctly identified. Using immunopeptideomics data, predicted off-target peptides that are also observed in mass spectrometry experiments can further indicate that the predicted off-targets may be genuine off-targets.

[0125] In step 106, similar peptides are identified. Similar peptides can be identified based on their presence in an organism (e.g., humans) and similarity to the target peptide at positions important for antigen recognition molecule binding. Similar peptides can be selected such that they are all the same length as the target peptide. Similar peptides can be identical to the target peptide, but from different protein sources than the target peptide. However, in some instances, the target peptides may have already been screened to ensure that non-unique candidate target peptides (candidate target peptides that share the same sequence with peptides elsewhere in the proteome) are not selected as targets.

[0126] A working list of peptide sequences can be collected to include peptide sequences in an organism (e.g., human). In some embodiments, the working list can be collected based on standard human protein sequences in a medical research database (such as UniprotKB, as a non-limiting example). The working list of peptide sequences can be filtered by comparing the peptide sequences in the database with the key positions of the target peptides identified in step 104 and retaining only peptide sequences that have identical amino acids to the target peptides at a threshold number of key positions. The threshold number of positions can be two positions, three positions, four positions, or five positions. In embodiments where the specific antigen recognition molecule is not known (i.e., the PIGSPRED method 100 does not rely on the antigen recognition molecule structure), three positions are preferred. Although there may be antigen recognition molecules that bind at only two positions, such antigen recognition molecules would likely have an impractically large number of off-target peptides. As the threshold number of positions is increased, off-targets may be underestimated.

[0127] The working list of peptides can be further filtered so that only peptides present in essential normal tissues or essential cell types are retained. Disruption of cellular functions of essential normal tissues or essential cell types is more likely to cause more severe off-target effects and therefore poses a greater risk compared to disruption of cellular functions of non-essential tissues. In some embodiments, a medical research database that includes gene expression in healthy donors can be utilized. Gene Tissue Expression Database or GTEx is a non-limiting example of a suitable database (see gtexportal.org.). For example, GTEx version 6 (v6) is based on genome build GRCh38 and contains gene expression data across 51 tissue types from 549 healthy donors. GTEx version 8 (v8) is based on genome build GRCh38 and contains gene expression data across 54 tissue types from 948 healthy donors; it is understood that more donor and tissue samples are being sequenced in newer releases. In some embodiments, alternative databases and / or different versions of the GTEx database can be utilized to determine potential high-risk off-target peptides in step 106, which may identify additional or fewer potential high-risk peptides as compared to the specific examples disclosed herein, as will be appreciated by those of skill in the art.

[0128] In some embodiments, gene expression values ​​are queried only in tissue types that are considered essential. Tissue types that may be considered non-essential may be those that may be sacrificed to save the patient's life, such as breast, ovary, testis, etc. In some embodiments, for each gene, the 95th percentile expression value is calculated in each essential tissue type, where the 95th percentile expression value means that 95% of all gene expression measurements in the tissue type are at or below that value. The maximum of the 95th percentile expression value across all essential tissue types is then calculated. If the maximum of the 95th percentile expression value of a gene is above 0.5 transcripts per millicon (TPM) (which may be adjusted based on transcription noise), it can be assumed that any similar peptides derived from the aforementioned gene would be potential high-risk off-target peptides. High TPM means high expression and vice versa. To calculate the TPM, the read counts for the genes are divided by the length of each gene in kilobases, thereby obtaining the reads per kilobase (RPK); then, the total RPK values ​​in the sample are counted and divided by 1,000,000, thereby obtaining a "parts per million" factor; finally, each RPK value is divided by the "parts per million" factor, thereby obtaining the TPM for each gene. See, for example, Li et al. (2010) Bioinformatics, 26(4): 493-500; Varabyou et al. (2021) Genome Res. 31(2): 301-308. The working list of peptide sequences can be filtered to remove peptides that are not identified as potentially high risk. In some embodiments, an alternative database to the GTEx database, and / or a different version of the GTEx database, can be utilized to determine potential high risk off-target peptides in step 106, which may identify additional or fewer potential high risk peptides compared to the specific examples disclosed herein, as will be appreciated by those skilled in the art.

[0129] In some embodiments, the peptides filtered from the working list of peptide sequences in step 106 may be tracked individually as a list of potential secondary target peptides. The secondary target peptides may be derived from secondary target genes different from the target genes of the target peptides, and cross-reactivity with the secondary target peptides may be therapeutically advantageous. If the potential secondary peptides are identical or highly similar to the target peptides, they are likely to cross-react with antigen recognition molecules that target the target peptides. Further testing (e.g., antigen recognition molecule screening) may be performed on one or more of the potential secondary target peptides as needed to assess whether cross-reaction with the potential secondary target peptides may provide advantageous therapeutic effects. Similar peptides (such as cancer-specific genes or peptides) from disease-associated genes or disease-associated antigens that have been identified as cancer-specific or tumor-specific antigens or that are derived from proteins that have been identified as cancer-specific or tumor-specific antigens may show advantageous cross-reactivity and thus may be secondary target peptides. As illustrated in Example 1, for example, cancer / testis (CT) antigens that are not expressed in normal essential tissues may be tracked if they are similar to the target peptides, especially if the target peptides are also CT antigens. Target peptides that share sequence identity with multiple suitable antigens for therapeutic targeting (eg, identical peptides or highly similar peptides) may be preferred.

[0130] In step 108, the working list of peptide sequences can be further filtered to remove peptides that do not bind to MHC molecules, and a degree of similarity (DoS) score can be calculated for the peptides remaining in the working list. The binding affinity of the peptides in the peptide sequence list can be determined by utilizing a computer-based binding affinity model. Peptides that are not predicted to bind to MHC molecules are removed from the working list of peptide sequences. In some embodiments, peptides with percentile rank values ​​of 2.0 or less are considered binders. Optionally, the list of potential secondary target peptides, if tracked, can be similarly filtered in the same manner as described for the working list of similar peptides in step 108 to remove peptides that do not bind to MHC molecules. Also, optional, a degree of similarity (DoS) score can be calculated for the potential secondary target peptides.

[0131] The DoS score can be based at least in part on the number of amino acids of each peptide in the working list that are identical to amino acids at corresponding positions of the target peptide identified in step 104 as important for binding to the antigen recognition molecule. In some embodiments, the DoS score considers only positions identified in step 104 as important for binding to the antigen recognition molecule. Alternatively, the DoS may consider at least some positions that were not identified in step 104 as important for binding to the antigen recognition molecule. For example, in some embodiments, the DoS score can be based on the number of amino acids identified (in step 116) as not involved in binding to the MHC molecule, or the DoS score can be based on the number of identical amino acids at every amino acid position of each peptide.

[0132] 2 shows a flow diagram of an embodiment of a method 200 for calculating expected off-target toxicity associated with an MHC-target peptide complex. Method 200 is preferably performed by a computing system, engine, module, device, and / or network. Method 200 preferably has instructions stored on a non-transitory computer readable medium.

[0133] In step 202, an MHC-target peptide complex is provided as an input.

[0134] In step 204, peptide positions (amino acids) important for antigen recognition molecule binding are identified. In a preferred embodiment, the peptide positions are identified according to method 104 shown in Figure 1B. In some embodiments, the peptide positions are identified by analyzing experimental structures of MHC-target peptide complexes with specific antigen recognition molecules.

[0135] In step 206, similar peptides can be identified. Similar peptides can be identified based on their presence in an organism (e.g., humans), similarity to the target peptide at positions important for antigen recognition molecule binding, and ability to bind to MHC molecules in MHC-target peptide complexes. Similar peptides can be selected such that they are all the same length as the target peptide. Similar peptides can be identical to the target peptide, but from different protein sources than the target peptide. However, in some instances, the target peptides may have already been screened so that non-unique candidate target peptides (candidate peptides that share the same sequence with peptides elsewhere in the proteome) are not selected as targets.

[0136] A working list of peptide sequences can be collected to include peptide sequences in an organism (e.g., human). The working list can be represented by a suitable data structure, such as a list, a matrix, and / or other suitable data structures understood by those skilled in the art. In some embodiments, the working list can be collected based on standard human protein sequences in medical research databases (such as, as a non-limiting example, UniprotKB). In some embodiments, step 206 can include receiving a peptide sequence of an organism (e.g., human) as input and utilizing the peptide sequence to generate a working list of peptide sequences.

[0137] The working list of peptide sequences can be filtered by comparing the peptide sequences in the database with the critical positions of the target peptides identified in step 204 and retaining only peptide sequences that have identical amino acids of the target peptides at or above a critical position number threshold. The position number threshold can be two positions, three positions, four positions, or five positions. In embodiments where the specific antigen recognition molecule is not known (i.e., the PIGSPRED method 100 does not depend on the antigen recognition molecule structure), three or four positions are preferred, and three positions are most preferred. The position number threshold is preferably determined based on the number of positions of binding for an antigen recognition molecule that is practical in a clinical setting. Although there may be antigen recognition molecules that bind at only two positions, such antigen recognition molecules are likely to have an impractically large number of off-target peptides, which may result in adverse side effects in a clinical setting, and therefore such antigen recognition molecules are not practical in a clinical setting. As the position number threshold is increased, off-targets for antigen recognition molecules that bind to positions fewer than the position number threshold may be underestimated; therefore, the expected off-target toxicity results calculated in step 210 may not be accurate for antigen recognition molecules that bind to positions fewer than the position number threshold.

[0138] In step 208, the expression of similar peptides in non-target cells is analyzed. The risk associated with antigen recognition molecule binding to each peptide of the similar peptides identified in step 206 is higher when each of the aforementioned peptides is present in essential normal tissues or essential cell types. Disruption of the cell function of essential normal tissues or essential cell types is more likely to cause off-target effects to be more severe, and therefore brings about a greater risk compared to disruption of the cell function of non-essential tissues. The working list of peptides can be further filtered so that only peptides present in essential normal tissues or essential cell types are retained.

[0139] In some embodiments, a medical research database containing gene expression in healthy donors can be utilized. The Gene Tissue Expression Database or GTEx is a non-limiting example of a suitable database (see gtexportal.org.). For example, GTEx version 6 (v6) is based on genome build GRCh38 and contains gene expression data across 51 tissue types from 549 healthy donors. GTEx version 8 (v8) is based on genome build GRCh38 and contains gene expression data across 54 tissue types from 948 healthy donors; it is understood that additional donor and tissue samples have been sequenced in newer releases. In some embodiments, gene expression values ​​are queried only in tissue types that are considered essential. Tissue types that may be considered non-essential may be those that can be sacrificed to save the patient's life, such as breast, ovary, testis, etc. In some embodiments, for each gene, the 95th percentile expression value in each essential tissue type is calculated, and then the maximum of the aforementioned values ​​across all essential tissue types is calculated. If the maximum expression value of a gene exceeds 0.5 TPM (which may be adjusted based on transcriptional noise), it can be assumed that any similar peptides derived from the aforementioned gene will be potential high-risk off-target peptides. The working list of peptide sequences can be filtered to remove peptides that are not identified as potentially high-risk. However, in some embodiments, peptides that are not identified as potentially high-risk may not need to be discarded. As discussed above, such peptides may form "secondary target peptides". Therefore, another working list containing potential secondary target peptides can be prepared, and further tests (e.g., antigen recognition molecule screening) can be optionally performed on one or more potential secondary target peptides to assess whether cross-reactivity with the aforementioned peptides can provide beneficial therapeutic effects.In some embodiments, alternative databases and / or different versions of the GTEx database can be utilized to determine potential high-risk off-target peptides in step 208, which may identify additional or fewer potential high-risk peptides as compared to the specific examples disclosed herein, as will be appreciated by those of skill in the art.

[0140] In step 210, the expected off-target toxicity associated with the MHC-target peptide complex is calculated. Once potential high-risk off-target peptides are identified, a degree of similarity (DoS) score is calculated for each of the high-risk off-target peptides. The DoS score quantifies the homology between the high-risk off-target peptide and the target peptide. DoS represents the number of identical amino acids at the same position or Hamming distance between the peptides. Using immunopeptideomics data, off-target peptides observed in mass spectrometry experiments can provide further evidence for potential off-targets.

[0141] The working list of peptide sequences can be filtered to remove peptides with a DoS score below a given threshold, and the number of peptide sequences remaining in the working list can be used as a metric indicative of off-target toxicity of the MHC-target peptide complex that was input in step 202 of the method 200 shown in FIG.

[0142] FIG. 3 shows a flow diagram of an embodiment of a method 300 for prioritizing potential target peptides to reduce off-target toxicity.

[0143] In step 302, two or more potential peptides are selected among the disease-associated peptides predicted to bind to the MHC molecule. Disease-specific MHC-target peptide complexes can be identified by identifying genes that are specifically expressed in diseased tissues. In some embodiments, disease-specific MHC-target peptide complexes can be identified based on medical databases that contain human gene sequences, such as The Cancer Genome Atlas (TCGA) and the Genome Tissue Expression Database (GTEx). In some embodiments, genes that are expressed in cancer types with a 75th percentile TPM value of more than 2 and have negligible expression in all essential normal tissues or essential cell types in GTEx are considered as cancer-specific genes. Standard protein sequences corresponding to cancer-specific genes can be derived from medical research databases that contain standard human protein sequences, such as the UniProtKB database.

[0144] The derived protein sequence can be used to predict potential 8-25 mer peptide sequences predicted to bind to the MHC of interest. When the MHC molecule is a class I MHC molecule, the predicted or detected peptides in step 302 can be 8-12 amino acids in length. When the MHC molecule is a class II MHC molecule, the predicted or detected peptides in step 302 can be 13-17 amino acids in length. Predictions can be made using binding affinity calculation tools such as the NetMHCpan tool. In some embodiments, a peptide is considered a binder if its predicted binding affinity is 500 nM or less; if it is 50 nM or less, it is considered a strong binder. In some embodiments, a peptide is considered a binder if its predicted percentile rank is 2 or less; if it is 0.5 or less, it is considered a strong binder.

[0145] The number of off-target peptides associated with each of the potential target peptides can be estimated at step 304. In some embodiments, the number of off-target peptides can be estimated using the steps of methods 100, 200 in FIG. 1A and / or FIG.

[0146] In step 306, the potential target peptides are ranked based at least in part on the number of off-target peptides associated with each of the potential target peptides. In some embodiments, the number of potential off-targets estimated in step 304 is representative of the likelihood of off-target toxicity that may be associated with the MHC-target peptide complex, and thus, this number can be used to rank the list of potential target peptides and prioritize such peptides for development of therapeutic agents.

[0147] After target selection and generation of therapeutic molecules that bind to the target, the off-targets predicted by PIGSPRED play an important role in experimental screening of therapeutic molecules for those that do not bind to the off-targets. Most specific therapeutic molecules can be selected for further development.

[0148] In some embodiments, the method 300 depicted in Figure 3 can be performed by repeating the method 200 in Figure 2, providing different target MHC-target peptide complexes as input at step 202, and obtaining a metric of off-target toxicity for each of the input MHC-target peptide complexes at step 210. The metrics of off-target toxicity can be compared to rank the MHC-target peptide complexes. MHC-target peptide complexes ranked with lower off-target toxicity can be selected for further analysis, can be used for generation of antigen recognition molecules, and / or can be used to test antigen recognition molecules.

[0149] 4 shows a block diagram of an exemplary embodiment of a computational system 400 configured to output a list of off-target peptides and / or a metric of off-target toxicity. The system 400 includes an exemplary embodiment of a PIGSPRED engine 410. The PIGSPRED engine 410 is configured to receive as input a computational representation of a target peptide presented in an MHC-target peptide complex 402. In the illustrated embodiment, the PIGSPRED engine 410 provides as output a list of off-target peptides 426 and / or a metric of off-target toxicity 428.

[0150] An illustrated embodiment of the PIGSPRED engine 410 includes a peptide location identification module 412, a working peptide list builder 414, a binding affinity filter 416, a risk severity assessment module 418, a risk severity filter 420, a risk likelihood assessment module 422, and a risk likelihood filter 424. The PIGSPRED engine 410 includes one or more processors and a non-transitory computer readable medium having instructions that, when executed by said one or more processors, cause the PIGSPRED engine 410 to perform functions associated with each of the above features 412, 414, 416, 418, 420, 422, 424. In some embodiments, the PIGSPRED engine can be configured to provide as output intermediate values ​​determined by the features 412, 414, 416, 418, 420, 422, 424 of the PIGSPRED engine 410. The features 412, 414, 416, 418, 420, 422, 424 are shown individually in a particular order for illustrative purposes. As will be understood by one of ordinary skill in the relevant art, the features 412, 414, 416, 418, 420, 422, 424 may be combined, permuted, or otherwise implemented in numerous configurations to achieve the disclosed functionality with respect to the illustrated embodiment of the PIGSPRED engine 410.

[0151] The illustrated system 400 further includes a binding affinity calculation engine 404, a standard protein sequence database 406, and a tissue expression database 408 in communication with a PIGSPRED engine 410. In a preferred embodiment, the binding affinity calculation engine 404, the standard protein sequence database 406, and the tissue expression database 408 are separate from the PIGSPRED engine 410, which is in communication with each other. Alternatively, one or more of the accessory features 404, 406, 408, 410 can be integrated into the PIGSPRED engine 400, as will be appreciated by those skilled in the relevant art. The binding affinity calculation engine 404 includes a computer-based model for calculating the binding affinity of peptides to MHC molecules. In some embodiments, the binding affinity calculation engine includes a commercially available tool, such as NetMHCpan or a similar product. The standard protein sequence database 406 includes information that may be derived from healthy tissue protein sequences. In some embodiments, the standard protein sequence database 406 includes the UniProtKB database or a similar medical database. As a non-limiting example, tissue expression database 408 can include gene expression data across approximately 30 tissue types from approximately 549 healthy donors; in another non-limiting example, tissue expression database 408 can include gene expression data across approximately 54 tissue types from approximately 948 healthy donors. In some embodiments, tissue expression database 408 comprises a gene tissue expression database, GTEx, or a similar medical database.

[0152] The peptide location identification module 412 of the PIGSPRED engine 410 receives as input a computational representation of the MHC-target peptide complex 402 and is capable of identifying positions important for binding to an antigen recognition molecule.

[0153] In a preferred embodiment, peptide location identification module 412 is configured to perform the following steps: (a) receiving as input a computational representation of a target peptide in an MHC-target peptide complex; (b) determining the binding affinity of the target peptide to the MHC molecule; (c) generating sequences of a plurality of mutated peptides, each associated with a mutation at a respective amino acid position of the target peptide; (d) determining the binding affinity of each mutated peptide of the plurality of mutated peptides to the MHC molecule; (e) predicting the amino acid locations involved in interaction with an antigen recognition molecule that recognizes the MHC-target peptide complex based in part on a comparison of the binding affinity of each mutated peptide to the binding affinity of the target peptide; and (f) providing as output a representation of the amino acid locations likely to be involved in interaction with the antigen recognition molecule.

[0154] Steps (b) and (d) are performed by utilizing a binding affinity calculation engine 404. In some embodiments, in step (e), a binding affinity threshold for the mutated peptide can be used to determine whether the amino acid position associated with the mutation is involved in the interaction with an antigen recognition molecule. In some embodiments, the predicted percentile rank is utilized to quantify the binding affinity, such that if the rank of the target peptide is 0.5 or less, the binding affinity threshold for the mutated peptide is 1.0; if the rank of the target peptide is greater than 0.5 and less than or equal to 2.0, the binding affinity threshold for the mutated peptide is 2.0; if the rank of the target peptide is greater than 2.0, the binding affinity threshold for the mutated peptide is 4.0. The mutated position is determined as the position that has lost binding affinity to the MHC molecule. The percentile rank is inversely proportional to the binding affinity, and therefore, when the binding affinity threshold for the mutated peptide is based on the percentile rank, a percentile rank below the threshold indicates that the amino acid position associated with the mutation is involved in the interaction with an antigen recognition molecule.

[0155] The working list builder 414 is configured to generate a working list of peptides such that within the total pool of predicted or detected preferred length peptides, the peptides listed in the working list (i) are located at positions corresponding to positions in the target peptide involved in the interaction with the antigen recognition molecule, and (ii) each contain at least two amino acids that are identical to corresponding amino acids in said target peptide. In some embodiments, the working peptide list builder 414 can derive the total pool of predicted or detected preferred length peptides from the standard protein sequence database 406. When the MHC molecule is a class I MHC molecule, the preferred length peptides can be 8-12 amino acids long. The preferred length peptides can be the same length as the target peptide. When the MHC molecule is a class II MHC molecule, the preferred length peptides can be 13-17 amino acids long. The preferred length peptides can be the same length as the target peptide.

[0156] The binding affinity filter 414 is configured to determine the binding affinity of each of the peptides listed in the working list to the MHC molecule and filter the working list to include only peptides having calculated binding affinities to the MHC molecule that show higher binding likelihood compared to a threshold. The binding affinity filter 416 can utilize the binding affinity calculation engine 404 to determine the binding affinity of each of the peptides in the working list to the MHC molecule. In some embodiments, peptides with a percentile rank value of 2.0 or less are considered binders by the binding affinity filter 414, and the percentile rank value is used as the threshold. The threshold used by the binding affinity filter 414 does not have to be equal to the threshold used by the location identification module 412.

[0157] The risk severity assessment module 418 is configured to estimate the number of peptides expressed in essential normal tissues or essential cell types in the working list of said off-target peptides. Disruption of the cellular function of essential normal tissues or essential cell types is more likely to cause more severe off-target effects and therefore poses a greater risk compared to disruption of the cellular function of non-essential tissues. The risk severity assessment module 418 can utilize the tissue expression database 408 to determine which peptides in the working list are expressed in essential normal tissues or essential cell types. In some embodiments, the gene expression values ​​are queried only in tissue types that are considered essential. Tissue types that may be considered non-essential may be those that may be sacrificed to save the patient's life, such as breast, ovary, testis, etc. In some embodiments, for each gene, the 95th percentile expression value in each essential tissue type is calculated, and then the maximum of said values ​​across all essential tissue types is calculated. If the maximum expression value of a gene is above 0.5 TPM (which may be adjusted based on transcriptional noise), it can be assumed that any similar peptides derived from said gene would be potential high-risk off-target peptides. A similar peptide may be identical to the target peptide, but from a different protein source than the target peptide. However, in some instances, the target peptide may have already been screened so that non-unique candidate target peptides (candidate target peptides that share an identical sequence with peptides elsewhere in the proteome) are not selected as targets.

[0158] The risk severity filter 420 filters the working list to remove peptides that were not identified as potentially high risk by the risk probability assessment module 418 .

[0159] The risk likelihood assessment module 422 calculates a degree of similarity (DoS) score for the peptides in the working list of peptides. The DoS score is based at least in part on the number of amino acids identical to amino acids at corresponding positions of the target peptide, which amino acids of said target peptide are involved in interaction with the antigen recognition molecule. In some embodiments, the DoS score considers only positions identified by the peptide position identification module 412 as important for binding to the antigen recognition molecule. Alternatively, the DoS may consider at least some positions that were not identified as important for binding to the antigen recognition molecule. For example, in some embodiments, the DoS score can be based on the number of amino acids that are not involved in binding to the MHC molecule, or the DoS score can be based on the number of identical amino acids at every amino acid position of the respective peptide. A larger DoS score is utilized as a predictor of the likelihood that an antigen recognition molecule that binds to an MHC-target peptide complex will also bind to a peptide in the working list of peptides. Thus, a larger likelihood of off-target peptide binding to an antigen recognition molecule reflects a larger risk likelihood. In some embodiments, for off-target peptides for which immunopeptidemics data are available, off-target peptides that have been observed in mass spectrometry experiments can provide further evidence of potential off-target similarity. For example, similar peptides observed in an immunopeptideme (by mass spectrometry) can be identified as off-target peptides based on having a DoS score that meets a DoS threshold that may be lower than a DoS threshold that would be used to identify similar peptides as off-target peptides when not otherwise observed in the immunopeptideme (e.g., in some implementations, a DoS score of 5 or more can qualify a similar peptide as an off-target peptide when observed in the immunopeptideme, whereas a DoS score of 6 or more can qualify a similar peptide as an off-target peptide when not otherwise observed in the immunopeptideme).

[0160] A risk probability filter 424 filters the working list to include only peptides with DoS scores above a threshold.

[0161] In the illustrated embodiment, the resulting working list represents a list of off-target peptides associated with the input MHC-target peptide complexes 402. In the illustrated embodiment, the PIGSPRED engine 410 provides as an output the resulting working list of off-target peptides.

[0162] In an illustrated embodiment, the total number of off-target peptides in the list of off-target peptides 426 is utilized as a metric of off-target toxicity 428, which may be provided as a second output to the PIGSPRED engine 410. In some embodiments, the metric of off-target toxicity 428 may be based at least in part on other factors, such as the DoS score of the off-target peptide (e.g., the average DoS score of all off-target peptides), a numerical quantification of the degree of "essentiality" of the tissue and / or cell type associated with the off-target peptide, protein amino acid sequence alignment, and / or a probability score for the off-target likelihood. Thus, as will be appreciated by one of skill in the relevant art informed by the teachings herein, the toxicity metric 428 may be a mathematical function of one or more factors related to the off-target peptide and the potential risk associated with the off-target peptide.

[0163] In the illustrated embodiment, the PIGSPRED engine further provides a list of verified off-target peptides 430 that includes peptides for which immunopeptidomic data is available and that have been used to verify the similarity of the off-target peptide to the target peptide.

[0164] In some embodiments, the PIGSPRED engine can be configured to output intermediate process results, including but not limited to binding affinities, DoS scores, and working lists at various filtering stages.

[0165] 5 shows a block diagram of an exemplary embodiment of a computational system 500 configured to output a list of low-risk peptide targets 558, a list of off-target peptides for potential targets 526, and / or metrics of off-target toxicity for potential targets 528. System 500 includes an exemplary embodiment of a target ranking engine 550.

[0166] The illustrated target ranking engine 550 includes a potential target selection module 552, a PIGSPRED module 510, and a target ranking module 554. The target ranking engine 550 includes one or more processors and a non-transitory computer-readable medium having instructions that, when executed by the one or more processors, cause the target ranking engine 550 to perform functions associated with each of the features 552, 510, 554 described above. In some embodiments, the target ranking engine 550 can be configured to provide as an output an intermediate value determined by the features 552, 510, 554 of the target ranking engine 550. The features 552, 510, 554 are shown individually in a particular order for illustration purposes. As will be appreciated by those skilled in the relevant art, the features 552, 510, 554 can be combined, permuted, or otherwise implemented in numerous configurations to achieve the disclosed functionality with respect to the illustrated embodiment of the target ranking engine 550.

[0167] The illustrated system 500 further includes a binding affinity calculation engine 504, a disease expression database 556, a standard protein sequence database 506, and a tissue expression database 508. The binding affinity calculation engine 504, the standard protein sequence database 506, and the tissue expression database 508 can be configured similarly to the corresponding features 404, 406, 408 in FIG. 4. The disease expression database can include information for identifying disease-specific MHC-target peptide complexes. In some examples, the disease expression database 556 can include a TCGA database that includes human gene sequences associated with cancer. In a preferred embodiment, the binding affinity calculation engine 504, the disease expression database 556, the standard protein sequence database 506, and the tissue expression database 508 are separate from the target ranking engine 550, which is in communication with each other. Alternatively, one or more of the accessory features 504, 556, 506, 508 can be integrated into the target ranking engine 550, as will be understood by those skilled in the relevant art.

[0168] The potential target selection module 552 is configured to communicate with a database of disease manifestations 556. The potential target selection module 552 is configured to determine a plurality of potential target peptides and associated MHC-target peptide complexes for each of said potential target peptides. The potential target selection module 552 can utilize the database of disease manifestations 556 to identify potential target peptides based on the type of disease under investigation, the prevalence of the potential target peptides, etc. The potential target selection module 552 can determine the binding affinity of the potential target peptides to one or more MHC molecules. The resulting MHC-target peptide complexes can be provided as input to the PIGSPRED module 510.

[0169] In the illustrated embodiment, the PIGSPRED module 510 functions similarly to the PIGSPRED engine 410 in Figure 4, where the MHC-target peptide complexes provided by the potential target selection module 552 are each provided as input MHC-target peptide complexes 402 to the system 400 in Figure 4, and the resulting list of off-target peptides 426, metrics of off-target toxicity 428, and list of verified off-target peptides 430 are each provided as output for each MHC-target peptide complex. In the illustrated embodiment, the list of off-target peptides for potential targets 526, metrics of off-target toxicity for potential targets 528, and list of verified off-target peptides 530 are provided as output to the target ranking engine 550.

[0170] The target ranking module 554 is configured to rank the MHC-target peptide complexes according to risk. In some examples, the target ranking module 554 can provide a list of low risk targets 558 as an output to the target ranking engine 550.

[0171] 6 shows a block diagram of an exemplary embodiment of a target toxicity database 600 that includes a list of MHC-target peptide complexes, their respective associated off-target peptides, and associated risk metrics. In a preferred embodiment, the off-target list and risk metrics can be determined for each MHC-target peptide complex using the system 400 shown in FIG. 4. In a preferred embodiment, the MHC-target peptide complexes can be ranked using the system shown in FIG. 5.

[0172] FIG. 7 illustrates a block diagram of an embodiment of a computing device 700. As shown, the computing device 700 may include one or more processors 710, an I / O device 720, an operating system ("OS") 740, a database 750, and a memory 730 including a program 760. In a preferred embodiment, the PIGSPRED engine 410 illustrated in FIG. 4 is implemented by the computing device 700. In another preferred embodiment, the target ranking engine 550 may be implemented by the computing device 700. In these preferred embodiments, instructions are stored in the memory 730, and the aforementioned instructions are executable by the processor 710 to perform the functions of the respective features 412, 414, 416, 418, 420, 422, 424, 552, 510, 554 of the respective engines 410, 510. In these preferred embodiments, the I / O device 720 is configured to communicate with the respective associated features 404, 406, 408, 504, 556, 506, 508 of the respective systems 400, 500 shown in FIGS.

[0173] In some embodiments, the computing device 700 may include some or all of the features of the PIGSPRED engine 410 shown in FIG. 4 and / or the target ranking engine 550 shown in FIG. 5, as well as additional features that may be implemented as instructions in the memory 730. In one exemplary embodiment, the instructions, when executed by the processor 710, cause the computing device 700 to provide as input a computational representation of an antigen recognition molecule, said antigen recognition molecule being capable of binding to an MHC-target peptide complex; determine the binding affinity of said antigen recognition molecule to a plurality of MHC-peptide complexes each including each likely off-target peptide from a working list and said MHC molecule; and filter said working list to include only off-target peptides that are likely to have binding affinity for the antigen recognition molecule, each MHC-peptide complex. An exemplary embodiment may be utilized to screen for specific therapeutic antigen recognition molecules. Assuming there is a list of off-targets that are independent of the use of the antigen recognition molecule, this exemplary embodiment finds off-targets for the specific antigen recognition molecule by removing off-targets from the independent list that do not bind to the antigen recognition molecule.

[0174] In another exemplary embodiment, the instructions, when executed by the processor 710, cause the computing device 700 to provide as input off-target peptide expression in essential normal tissues or essential cell types of a particular patient; and provide as output an indication of off-target effects in said patient. Exemplary embodiments may be utilized to screen patients for clinical trials or other treatments.

[0175] The computing device 700 may be a single server or may be configured as a distributed computer system including multiple servers or computers that interoperate to perform one or more of the processes and functions associated with the disclosed embodiments. In some embodiments, the computing device 700 may further include a peripheral interface in communication with the processor 710, a transceiver, a mobile network interface, a bus configured to facilitate communication between various components of the computing device 700, and a power supply configured to provide power to one or more components of the computing device 700. The peripheral interface may include hardware, firmware, and / or software capable of communicating with various peripheral devices, such as media drives (e.g., magnetic disk, solid-state, or optical disk drives), other processing devices, or any other input source used in connection with the techniques of the present invention. In some embodiments, the peripheral interface may include a serial port, a parallel port, a general purpose input / output (GPIO) port, a game port, a universal serial bus (USB), a micro USB port, a high definition multimedia (HDMI) port, a video port, an audio port, a Bluetooth port, an NFC port, another similar communication interface, or any combination thereof.

[0176] In some embodiments, the transceiver may be configured to communicate with compatible devices and ID tags when they are within a predetermined range. The transceiver may be compatible with one or more of the following: RFID, NFC, Bluetooth, Low Energy Bluetooth (BLE), WiFi™, ZigBee, ABC protocol, or similar technologies.

[0177] The mobile network interface may access a cellular network, the Internet, or another wide area network. In some embodiments, the mobile network interface may include hardware, firmware, and / or software that enables the processor 710 to communicate with other devices over local or wide area, private or public, wired or wireless networks known in the art. The power source may be configured to provide alternating current (AC) or direct current (DC) to the components as appropriate.

[0178] The processor 710 may include one or more of a microprocessor, microcontroller, digital signal processor, or co-processor, or the like, or combinations thereof, capable of executing stored instructions and manipulating on data storage. The memory 730 may include one or more suitable types of memory for sorting files (e.g., volatile or non-volatile memory, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic disk, optical disk, floppy disk, hard disk, removable cartridge, flash memory, redundant array of independent disks (RAID), and the like), in some implementations, including an operating system, application programs (e.g., including a web browser application, widget or gadget engine, or other application, as appropriate), executable instructions, and data. In one embodiment, the processing techniques described herein are implemented as a combination of executable instructions and data in the memory 730.

[0179] The processor 710 may be one or more known processing devices, such as the Pentium family of microprocessors manufactured by Intel™ or the Turion family of microprocessors manufactured by AMD™. The processor 710 may be comprised of a single core processor or a multi-core processor that simultaneously executes parallel processes. For example, the processor 710 may be a single core processor configured with virtual processing technology. In certain embodiments, the processor 710 may use a logical processor for simultaneously executing and controlling multiple processes. The processor 710 may implement virtual machine technology or other similar known technologies for providing the ability to, for example, execute, control, operate, manipulate, store, or perform multiple software processes, applications, programs, and the like. As will be appreciated by those skilled in the relevant art, other processor arrangement types may be implemented that provide the capabilities disclosed herein.

[0180] The computing device 700 may include one or more storage devices configured to store information used by the processor 710 (or other components) to perform certain functions related to the disclosed embodiments. In one example, the computing device 700 may include a memory 730 that includes instructions to enable the processor 710 to execute one or more applications, such as server applications, network communication processes, and any other type of application or software known to be available in computer systems. Alternatively, the instructions, application programs, etc. may be stored in an external storage device or available from the memory over a network. The one or more storage devices may be volatile or non-volatile, magnetic, semiconductor, tape, optical, removable, non-removable, or other types of storage devices or tangible computer-readable media.

[0181] In one embodiment, the computing device 700 may include a memory 730 that includes instructions that, when executed by the processor 710, perform one or more processes consistent with the functionality disclosed herein. Methods, systems, and products consistent with the disclosed embodiments are not limited to individual programs or computers configured to perform dedicated tasks. For example, the computing device 700 may include a memory 730 that may include one or more programs 760 for performing one or more functions of the disclosed embodiments. Furthermore, the processor 710 may execute one or more programs 760 that are located remotely from the computing device 700. For example, the computing device 700 may access one or more remote programs 760 that, when executed, perform functions related to the disclosed embodiments.

[0182] Memory 730 may include one or more memory devices that store data and instructions used to implement one or more features of the disclosed embodiments. Memory 730 may also include any combination of one or more databases controlled by a memory controller device (e.g., a server, etc.) or software, such as a document management system, a Microsoft™ SQL database, a SharePoint™ database, an Oracle™ database, a Sybase™ database, or other related databases. Memory 730 may include software components that, when executed by processor 710, cause one or more processes consistent with the disclosed embodiments to be performed. In some embodiments, memory 730 may include a database 750 for storing relevant data to enable computing device 700 to implement one or more of the processes and functionalities associated with the disclosed embodiments.

[0183] Computing device 700 may also be communicatively connected to one or more memory devices, such as a database (not shown), locally or over a network. The remote memory device may be configured to store information and accessed and / or managed by computing device 700. By way of example, the remote memory device may be a document management system, a Microsoft™ SQL database, a SharePoint™ database, an Oracle™ database, a Sybase™ database, or other related database. However, systems and methods consistent with disclosed embodiments are not limited to a particular database, or even to the use of a database.

[0184] Computing device 700 may also include one or more I / O devices 720, which may include one or more interfaces that receive signals or inputs from devices and provide signals or outputs to one or more devices thereby enabling data to be received and / or transmitted by computing device 700. For example, computing device 700 may include an interface component that may provide an interface to one or more input devices (such as one or more keyboards, a mouse device, a touch screen, a track pad, a track ball, a scroll wheel, a digital camera, a microphone, and a sensor) that enable computing device 700 to receive data from one or more users (e.g., via user device 130).

[0185] In an exemplary embodiment of the disclosed technology, computing device 700 may include any number of hardware and / or software applications executing to facilitate any of its operations. One or more I / O interfaces may be utilized to receive or collect data and / or user instructions from a wide variety of input devices. Received data may be processed by one or more computer processors and / or stored in one or more memory devices as desired in various implementations of the disclosed technology.

[0186] Although computing device 700 is described as one form for implementing the techniques described herein, other functionally equivalent techniques may be used, as will be appreciated by those skilled in the relevant art. For example, as is known in the art, some or all of the functionality implemented via executable instructions may also be implemented using firmware and / or hardware devices (such as application specific integrated circuits (ASICs), programmable logic arrays, state machines, etc.). Additionally, other implementations may include more or fewer components than those shown.

[0187] 8 illustrates a block diagram of an embodiment of a computing network 800 including a computing device 810, a server 820, memory storage 840, and a network 830 facilitating communication between each. In a preferred embodiment, the computing device 810 includes a computing device configured as shown in FIG. 7 and including the PIGSPRED engine 410 in FIG. 4 and / or the target ranking engine 550 in FIG. 5. In some embodiments, the computing device 810 communicates with some or all of the attached features 404, 406, 408, 504, 556, 506, 508 via the network 830, where the aforementioned attached features are hosted on the server 820 and accessible from the memory storage 840.

[0188] Network 830 may be any suitable type of network, including a separate connection over the Internet, such as a cellular network or a WiFi™ network. In some embodiments, network 830 may connect to terminals, services, and mobile devices using direct connections, such as Radio Frequency Identification (RFID), Near Field Communication (NFC), Bluetooth, Low Energy Bluetooth (BLE), WiFi, ZigBee, Ambient Backscatter Communication (ABC) protocol, USB, WAN, or LAN. Security concerns may dictate that one or more of these types of connections be encrypted or otherwise protected, since the information transmitted may be personal or confidential. However, in some embodiments, the information transmitted may be less personal, and thus a network connection may be selected that prioritizes convenience over security.

[0189] 9 illustrates cellular functions for exemplary embodiments presented herein. A cell 910 includes MHC-peptide complexes 920, 930 present on a surface 912 of the cell 910. The cell 910 includes a class I MHC-peptide complex 920 and a class II MHC-peptide complex 930. Each MHC-peptide complex 920, 930 includes a respective MHC molecule 922, 932 and a respective peptide 924, 934. Each peptide 924, 934 includes amino acids bound to the respective MHC molecule 922, 932 (shown as shaded objects) and amino acids not bound to the respective MHC molecule 922, 932 (shown as white objects). An antigen recognition molecule 940 includes a receptor 942 capable of binding to amino acids of a peptide in the MHC-target peptide complex that is not bound to an MHC molecule. Additionally, the antigen recognition molecule 940 may bind to an off-target peptide that has a similar amino acid composition (as compared to the MHC-target peptide complex) that is not bound to the respective MHC molecule in the MHC-peptide complex.

[0190] Various methods described herein may involve the identification (e.g., ranking / selection) of target peptides and / or the identification (e.g., ranking / selection) of off-target peptides of a given target peptide. Any of these methods may further comprise the step of synthesizing the identified target peptide and / or one or more off-target peptides. Methods of peptide synthesis are known in the art and include the methods described herein. Any of these methods may further comprise the step of loading the target peptide and / or one or more off-target peptides onto an MHC molecule or any suitable component thereof to form a pMHC complex, as described elsewhere herein. Any of these methods may further comprise the step of binding the target peptide-MHC complex and / or one or more off-target peptide-MHC complex to an antigen recognition molecule (e.g., an antibody, a TCR, or a CAR). For example, any of these methods may further comprise the step of screening the target peptide-MHC complex and / or one or more off-target peptide-MHC complex for binding to an antigen recognition molecule. Any of these methods may include incubating a target peptide and / or one or more off-target peptides with one or more cells (e.g., pulsing the cells with the peptides as described elsewhere herein).

[0191] In a further aspect, the present disclosure provides herein an off-target peptide identified using the method described herein.Therefore, the present disclosure also provides a library that comprises one or more of the off-target peptides identified using the method described herein.In some embodiments, the library of the present disclosure can also comprise a target peptide associated with the off-target peptide.In some embodiments, the library of the present disclosure can comprise an off-target peptide identified by analyzing the experimental structure of pHLA in complex with an antigen recognition molecule.

[0192] Exemplary peptide libraries of the present disclosure include the MAGEA4 230~239 In some embodiments, the MAGEA4 gene may include one or more off-target peptides identified for the target GVYDGREHTV (SEQ ID NO: 1). 230~239 Off-target peptides related to the target GVYDGREHTV (SEQ ID NO: 1) include those listed in Tables 1B, 2B, and 3B herein. In some embodiments, MAGEA4 230~239 The off-target peptide related to the target GVYDGREHTV (SEQ ID NO: 1) comprises any of the amino acid sequences of SEQ ID NOs: 2-7, 9, 11-28, and 73-74, or a pharma- ceutically acceptable salt thereof, or a fragment or derivative thereof. 230~239 The off-target peptide associated with the target GVYDGREHTV (SEQ ID NO: 1) consists essentially of any of the amino acid sequences of SEQ ID NOs: 2-7, 9, 11-28, and 73-74. 230~239 The off-target peptide related to the target GVYDGREHTV (SEQ ID NO: 1) consists of any one of the amino acid sequences of SEQ ID NOs: 2-7, 9, 11-28, and 73-74.

[0193] Another exemplary peptide library of the disclosure is the MAGEA4 286~294 In some embodiments, the MAGEA4 polypeptide may include one or more off-target peptides identified for the target KVLEHVVRV (SEQ ID NO: 48). 286~294 Off-target peptides related to the target KVLEHVVRV (SEQ ID NO: 48) include those listed in Table 4B herein. In some embodiments, MAGEA4 286~294 The off-target peptide related to the target KVLEHVVRV (SEQ ID NO: 48) comprises any of the amino acid sequences of SEQ ID NOs: 49-72, or a pharma- ceutically acceptable salt thereof, or a fragment or derivative thereof. 286~294The off-target peptide associated with the target KVLEHVVRV (SEQ ID NO: 48) consists essentially of the amino acid sequence of any of SEQ ID NOs: 49-72. 286~294 The off-target peptide related to the target KVLEHVVRV (SEQ ID NO: 48) consists of any one of the amino acid sequences of SEQ ID NOs: 49 to 72.

[0194] Another exemplary peptide library of the disclosure is the MAGEA3 168~176 In some embodiments, the MAGEA3 peptide may include one or more off-target peptides identified for the target EVDPIGHLY (SEQ ID NO: 29). 168~176 Off-target peptides related to the target EVDPIGHLY (SEQ ID NO: 29) include those listed in Table 5B herein. In some embodiments, MAGEA3 168~176 The off-target peptide related to the target EVDPIGHLY (SEQ ID NO: 29) comprises any of the amino acid sequences of SEQ ID NOs: 30-47, or a pharma- ceutically acceptable salt thereof, or a fragment or derivative thereof. 168~176 The off-target peptide associated with the target EVDPIGHLY (SEQ ID NO: 29) consists essentially of the amino acid sequence of any of SEQ ID NOs: 30-47. 168~176 The off-target peptide related to the target EVDPIGHLY (SEQ ID NO: 29) consists of any one of the amino acid sequences of SEQ ID NOs: 30 to 47.

[0195] Yet further exemplary peptide libraries of the present disclosure may include one or more peptides selected from: MAGEA4 230~239 Target GVYDGREHTV (SEQ ID NO: 1), SEQ ID NOs: 2-7, 9, 11-28, and 73-74; MAGEA4 286~294 Target KVLEHVVRV (SEQ ID NO: 48), SEQ ID NO: 49-72; MAGEA3 168~176Target EVDPIGHLY (SEQ ID NO: 29), and SEQ ID NOs: 30 to 47, or any combination thereof.

[0196] The peptide libraries of the present disclosure can include any number of peptides as desired. For example, the peptide libraries of the present disclosure can include at least 2 peptides, at least 3 peptides, at least 4 peptides, at least 5 peptides, at least 6 peptides, at least 7 peptides, at least 8 peptides, at least 9 peptides, at least 10 peptides, at least 20 peptides, at least 30 peptides, at least 40 peptides, at least 50 peptides, or about 2-5 peptides, about 2-10 peptides, about 5-15 peptides, about 10-20 peptides, about 10-30 peptides, about 12-25 peptides, about 20-30 peptides, about 25-50 peptides, about 40-80 peptides, about 50-100 peptides, etc.

[0197] In a further aspect, provided herein is a database that includes computational representations of off-target peptides identified using the methods described herein. The database can include information about each of the off-target peptides represented in any one of the exemplary libraries disclosed above. Thus, the database can also include computational representations of target peptides associated with the off-target peptides. The computer readable representations of peptides in the database can include peptide sequences or computer models of peptides. For example, the database can include features of database 600 shown in FIG. 6, including one or more lists of off-target peptides. Database 600 may or may not include associated MHC-target peptide complexes and may or may not include associated risk metrics.

[0198] The peptides of the present disclosure can be produced synthetically or hydrolyzed.Synthetically produced peptides can include randomly generated peptides, specifically designed peptides, and peptides in which at least some of the amino acid positions are conserved among several peptides, and the remaining positions are random.Alternatively, the peptides of the present disclosure can be produced by expression in a heterologous host cell.

[0199] In some embodiments, the peptides of the present disclosure can be synthesized, for example, by solid-phase synthesis. As such, the peptides can be immobilized on a solid support, such as, for example, beads. The peptides of the present disclosure can be synthesized by the Fmoc-polyamide method of solid-phase peptide synthesis. The N-amino group is transiently protected by the 9-fluorenylmethyloxycarbonyl (Fmoc) group. Repetitive cleavage of this highly base-labile protecting group is performed using 20% ​​piperidine in N,N-dimethylformamide. Side chain functional groups can be protected as butyl ethers (for serine, threonine, and tyrosine), butyl esters (for glutamic acid and aspartic acid), butyloxycarbonyl derivatives (for lysine and histidine), trityl derivatives (for cysteine), and 4-methoxy-2,3,6-trimethylbenzenesulfonyl derivatives (for arginine) of the aforementioned functional groups. When glutamine or asparagine is the C-terminal residue, a 4,4'-dimethoxybenzhydryl group is used for protection of the side chain amide functionality. The solid support is based on a polydimethyl-acrylamide polymer composed of the three monomers dimethylacrylamide (backbone monomer), bisacryloylethylenediamine (crosslinker), and acryloylsarcosine methyl ester (functionalizer). The peptide-resin cleavable linking agent used is an acid-labile 4-hydroxymethyl-phenoxyacetic acid derivative. All amino acid derivatives are added as preformed symmetrical anhydride derivatives, except for asparagine and glutamine, which are added using the reverse N,N-dicyclohexyl-carbodiimide / 1-hydroxybenzotriazole mediated coupling procedure. All coupling and deprotection reactions are monitored using ninhydrin, trinitrobenzenesulfonic acid, or isotonic acid test procedures. Upon completion of the synthesis, the peptides are cleaved from the resin support and simultaneously the side chain protecting groups are removed by treatment with 95% trifluoroacetic acid containing a 50% scavenger mixture. Commonly used scavengers include ethanedithiol, phenol, anisole, and water, and are precisely selected according to the amino acid composition of the peptide to be synthesized. Also, a combination of solid and liquid phase methods for synthesizing peptides is possible.

[0200] Trifluoroacetic acid is removed by evaporation in vacuo, followed by trituration with diethyl ether to give the crude peptide. Any scavenger present is removed by a brief extraction procedure in which the aqueous phase is lyophilized to give the scavenger-free crude peptide.

[0201] Purification may be achieved by techniques such as recrystallization, ion exchange chromatography, size exclusion chromatography, hydrophobic interaction chromatography, and reverse-phase high performance liquid chromatography (eg, using an acetonitrile / water gradient separation), or a combination of these.

[0202] Peptides may be analyzed using thin layer chromatography, electrophoresis, particularly capillary electrophoresis, solid phase extraction (CSPE), reversed-phase high performance liquid chromatography, amino acid analysis after acid hydrolysis, and fast electron impact (FAB) mass spectrometry, as well as MALDI and ESI-Q-TOF mass spectrometry.

[0203] Alternatively, the peptides can be produced by recombinant expression in heterologous host cells. Such methods typically involve the use of a vector that contains a nucleic acid sequence encoding the peptide to be expressed in vivo (e.g., in bacterial, yeast, insect, or mammalian cells), the polypeptide to be expressed.

[0204] In further embodiments, an in vitro cell-free system may be used. The peptides may be isolated and / or provided in substantially pure form. For example, the peptides may be provided in a form that is substantially free of other peptides or proteins.

[0205] In some embodiments, the peptides of the present disclosure may be about 8-25 amino acids in length. In some embodiments, the peptides of the present disclosure may be about 8-12 amino acids in length. For example, the peptides disclosed herein may be 8, 9, 10, 11, or 12 amino acids in length. In some embodiments, the peptides of the present disclosure may be about 13-17 amino acids in length. For example, the peptides disclosed herein may be 13, 14, 15, 16, or 17 amino acids in length.

[0206] The peptides of the present disclosure may include one or more chemical modifications. Non-limiting examples of chemical modifications include, for example, phosphorylation, acetylation, deamidation, acylation, amidination, pyridoxylation of lysine, reductive alkylation, trinitrobenzylation of amino groups with 2,4,6-trinitrobenzenesulfonic acid (TNBS), amide modification of carboxyl groups and sulfhydryl modification of cysteine ​​to cysteic acid by performic acid oxidation, formation of mercury derivatives, formation of disulfide mixtures with other thiol compounds, reaction with maleimides, carboxymethylation with iodoacetic acid or iodoacetamide, and carbamoylation with cyanate at alkaline pH. Chemical modifications may not correspond to those that may exist in vivo.

[0207] For example, the modification of arginyl residues, for example in proteins, can be based on the reaction of vicinal dicarbonyl compounds (such as phenylglyoxal, 2,3-butanedione, and 1,2-cyclohexanedione) to form adducts. Another example is the reaction of methylglyoxal with arginine residues. Cysteines can be modified without concomitant modification of other nucleophilic sites such as lysines and histidines. Disulfide bonds in proteins can also be selectively reduced. Disulfide bonds can be formed and oxidized during thermal treatment of biopharmaceuticals. Woodward's reagent K can be used to modify specific glutamic acid residues. N-(3-(dimethylamino)propyl)-N'-ethylcarbodiimide can be used to form intramolecular crosslinks between lysine and glutamic acid residues. For example, diethylpyrocarbonate and 4-hydroxy-2-nonenal can be used to modify histidyl residues in proteins. Reactions of lysine residues and other α-amino groups are useful, for example, in binding peptides to surfaces or protein / peptide crosslinking. Lysine is the attachment site of poly(ethylene) glycol and the primary modification site in protein glycosylation. Methionine residues in proteins can be modified with, for example, iodoacetamide, bromoethylamine, and chloramine T. Tetranitromethane and N-acetylimidazole can be used for the modification of tyrosyl residues. Cross-linking via the formation of dityrosine can be performed with hydrogen peroxide / copper ions. N-bromosuccinimide, 2-hydroxy-5-nitrobenzyl bromide, or 3-bromo-3-methyl-2-(2-nitrophenylmercapto)-3H-indole (BPNS-skatole) have been used for the modification of tryptophan in recent studies.

[0208] The peptides described herein may include one or more (e.g., one, two, three, or four) amino acid substitutions and / or insertions and / or deletions. Amino acid substitution means replacing an amino acid residue with a replacement amino acid residue at the same position. The inserted amino acid residue may be inserted at any position, and may be inserted such that some or all of the inserted amino acid residues are directly adjacent to each other, or may be inserted such that the inserted amino acid residues are not directly adjacent to another inserted amino acid residue. One or more (e.g., one, two, three, or four) amino acids may be substituted and / or inserted and / or deleted in any one of the sequences of SEQ ID NOs: 1-7, 9, and 11-74. Each substitution and / or insertion and / or deletion may occur at any position in any one of SEQ ID NOs: 1-7, 9, and 11-74.

[0209] In some embodiments, peptides of the disclosure may include additional amino acids (e.g., 1, 2, 3, or 4) at the C-terminus and / or N-terminus of any one of SEQ ID NOs: 1-7, 9, and 11-74. Peptides of the disclosure may include the amino acid sequence of any one of SEQ ID NOs: 1-7, 9, and 11-74, except for one or more (e.g., 1, 2, 3, or 4) amino acid substitutions, insertions, or deletions.

[0210] Amino acid substitutions may be conservative, meaning that the substituted amino acid has similar chemical properties as the original amino acid. For example, the following groups of amino acids share similar chemical properties such as size, charge, and polarity: Group 1 - Ala, Ser, Thr, Pro, Gly; Group 2 - Asp, Asn, Glu, Gln; Group 3 - His, Arg, Lys; Group 4 - Met, Leu, Ile, Val, Cys; Group 5 - Phe, Thy, Trp.

[0211] In another aspect, the present disclosure provides a complex of the peptide of the present disclosure and an MHC molecule (pMHC complex). Preferably, the peptide is bound to the peptide-receiving groove of the MHC molecule. In some embodiments, the peptide and the MHC molecule form a non-covalent complex. In other embodiments, the peptide and the MHC molecule may be covalently linked, for example, via a linker. Thus, the present disclosure also provides a library comprising one or more of the pMHC complexes described herein.

[0212] Exemplary pMHC complex libraries of the present disclosure include the MAGEA4 230~239 In some embodiments, the pMHC complex may include one or more pMHC complexes that include an off-target peptide related to the target GVYDGREHTV (SEQ ID NO: 1). 230~239 Off-target peptides related to the target GVYDGREHTV (SEQ ID NO: 1) include those listed in Tables 1B, 2B, and 3B herein. In some embodiments, MAGEA4 present in the pMHC complex 230~239 The off-target peptide associated with the target GVYDGREHTV (SEQ ID NO: 1) comprises any of the amino acid sequences of SEQ ID NOs: 2-7, 9, 11-28, and 73-74, or a pharma- ceutically acceptable salt thereof, or a fragment or derivative thereof. In some embodiments, the MAGEA4 peptide present in the pMHC complex 230~239 The off-target peptide associated with the target GVYDGREHTV (SEQ ID NO: 1) consists essentially of any of the amino acid sequences of SEQ ID NOs: 2-7, 9, 11-28, and 73-74. In some embodiments, the MAGEA4 peptide present in the pMHC complex 230~239 The off-target peptide related to the target GVYDGREHTV (SEQ ID NO: 1) consists of any one of the amino acid sequences of SEQ ID NOs: 2-7, 9, 11-28, and 73-74.

[0213] Another exemplary pMHC complex library of the present disclosure is the MAGEA4 286~294In some embodiments, the pMHC complex may include one or more pMHC complexes that include an off-target peptide related to the target KVLEHVVRV (SEQ ID NO: 48). 286~294 Off-target peptides related to the target KVLEHVVRV (SEQ ID NO: 48) include those listed in Table 4B herein. In some embodiments, MAGEA4 present in the pMHC complex 286~294 The off-target peptide associated with the target KVLEHVVRV (SEQ ID NO: 48) comprises any of the amino acid sequences of SEQ ID NOs: 49-72, or a pharma- ceutically acceptable salt thereof, or a fragment or derivative thereof. In some embodiments, the MAGEA4 peptide present in the pMHC complex 286~294 The off-target peptide associated with the target KVLEHVVRV (SEQ ID NO: 48) consists essentially of the amino acid sequence of any of SEQ ID NOs: 49-72. In some embodiments, the MAGEA4 peptide present in the pMHC complex 286~294 The off-target peptide related to the target KVLEHVVRV (SEQ ID NO: 48) consists of any one of the amino acid sequences of SEQ ID NOs: 49 to 72.

[0214] Another exemplary pMHC complex library of the present disclosure is the MAGEA3 168~176 In some embodiments, the MAGEA3 may comprise one or more pMHC complexes that contain an off-target peptide related to the target EVDPIGHLY (SEQ ID NO: 29). 168~176 Off-target peptides related to the target EVDPIGHLY (SEQ ID NO: 29) include those listed in Table 5B herein. In some embodiments, MAGEA3 present in the pMHC complex 168~176 The off-target peptide associated with the target EVDPIGHLY (SEQ ID NO: 29) comprises any of the amino acid sequences of SEQ ID NOs: 30-47, or a pharma- ceutically acceptable salt thereof, or a fragment or derivative thereof. In some embodiments, the MAGEA3 peptide present in the pMHC complex 168~176The off-target peptide associated with the target EVDPIGHLY (SEQ ID NO: 29) consists essentially of the amino acid sequence of any of SEQ ID NOs: 30-47. In some embodiments, the MAGEA3 peptide present in the pMHC complex 168~176 The off-target peptide related to the target EVDPIGHLY (SEQ ID NO: 29) consists of any one of the amino acid sequences of SEQ ID NOs: 30 to 47.

[0215] Yet further exemplary pMHC complex libraries of the present disclosure may include one or more pMHC complexes comprising a peptide selected from: MAGEA4 230~239 Target GVYDGREHTV (SEQ ID NO: 1), SEQ ID NOs: 2-7, 9, 11-28, and 73-74; MAGEA4 286~294 Target KVLEHVVRV (SEQ ID NO: 48), SEQ ID NO: 49-72; MAGEA3 168~176 Target EVDPIGHLY (SEQ ID NO: 29), and SEQ ID NOs: 30 to 47, or any combination thereof.

[0216] The pMHC complex libraries of the present disclosure may contain any number of multiple pMHC complexes as desired. For example, a pMHC complex library of the present disclosure may include at least 2 pMHC complexes, at least 3 pMHC complexes, at least 4 pMHC complexes, at least 5 pMHC complexes, at least 6 pMHC complexes, at least 7 pMHC complexes, at least 8 pMHC complexes, at least 9 pMHC complexes, at least 10 pMHC complexes, at least 20 pMHC complexes, at least 30 pMHC complexes, at least 40 pMHC complexes, at least 50 pMHC complexes, or about 2-5 pMHC complexes, about 2-10 pMHC complexes, about 5-15 pMHC complexes, about 10-20 pMHC complexes, about 10-30 pMHC complexes, about 12-25 pMHC complexes, about 20-30 pMHC complexes, about 25-50 pMHC complexes, about 40-80 pMHC complexes, about 50-100 pMHC complexes, etc.

[0217] MHC molecules used in the pMHC complexes described herein include naturally occurring full-length MHC molecules, as well as the individual chains of MHC molecules (e.g., MHC class I α (heavy) chain, β2-microglobulin, MHC class II α chain, and MHC class II β chain), the individual subunits of such chains of MHC (e.g., the α1, α2, and / or α3 subunits of the MHC class I α chain, the α1 and / or α2 subunits of the MHC class II α chain, the β1 and / or β2 subunits of the MHC class II β chain), and fragments, mutants, and various derivatives thereof (including fusion proteins, e.g., fusions with viral envelope proteins or fusogens), where such fragments, mutants, and derivatives retain the ability to display antigenic determinants for recognition by an antigen recognition molecule.

[0218] Naturally occurring MHC molecules are encoded by a group of genes on human chromosome 6 or mouse chromosome 17. MHC is also referred to as H-2 in mice and human leukocyte antigen (HLA) in humans. MHC class I molecules specifically bind to CD8 molecules expressed on cytotoxic T lymphocytes (CD8+ T cells), whereas MHC class II molecules specifically bind to CD4 molecules expressed on helper T lymphocytes (CD4+ T cells). MHC includes, but is not limited to, HLA specificities such as A (e.g., A1-A74), B (e.g., B1-B77), C (e.g., C1-C11), D (e.g., D1-D26), E, ​​G, DR (e.g., DR1-DR8), DQ (e.g., DQ1-DQ9), and DP (e.g., DP1-DP6). More preferably, the HLA specificities include A1, A2, A3, A11, A23, A24, A28, A30, A33, B7, B8, B35, B44, B53, B60, B62, DR1, DR2, DR3, DR4, DR7, DR8, and DR-11.

[0219] In some embodiments, the MHC molecule in the pMHC complex of the present disclosure is a human leukocyte antigen (HLA) molecule. The MHC molecule can be a human HLA molecule selected from the group consisting of HLA-A, HLA-B, HLA-C, HLA-E, HLA-F, and HLA-G. In some embodiments, the MHC class I or MHC II polypeptide can be derived from any functional human HLA-A, B, C, DR, or DQ molecule. Non-limiting examples of HLA-A alleles include A, B, C, DR, DQ, and DQ. * 01:01, A * 02:01, A * 02:02, A * 03:01, A * 11:01, A * 23:01, A * 24:02, A * 25:01, A * 26:01, A * 29:01, A * 29:02, A * 31:01, A * 32:01, A * 33:01, A * 34:01, A * 36:01, A * 43:01, A * 66:01, A * 68:01, A * 69:01, A * 74:01, and A * 80:01. Non-limiting examples of HLA-B alleles include B * 07:02.B * 08:01, B * 13:01, B * 14:01, B * 14:02, B * 15:01, B * 18:01, B * 18:02, B * 27:01, B * 27:02, B * 35:01, B * 35:02, B * 37:01, B * 38:01, B * 39:01, B *40:01, B * 41:01, B * 42:01, B * 44:02, B * 45:01, B * 46:01, B * 47:01, B * 48:01, B * 49:01, B * 50:01, B * 51:01, B * 52:01, B * 53:01, B * 54:01, B * 55:01, B * 55:02, B * 56:01, B * 57:01, B * 58:01, B * 59:01, B * 67:01, B * 73:01, B * 15:17, B * 81:01, B * 82:01, and B * 83:01. Non-limiting examples of HLA-C alleles include Cw * 01:01, Cw * 02:02, Cw * 03:03, Cw * 04:01, Cw * 05:01, Cw * 06:02, Cw * 07:01, Cw * 07:02, Cw * 08:02, Cw * 12:03, Cw * 14:01, Cw * 15:02, Cw * 16:01, Cw * 17:01. and Cw * 18:01. Non-limiting examples of HLA-DR alleles include DRB1 * 01:01, DRB1 * 01:03, DRB1 * 15:01, DRB1 * 15:02, DRB1* 16:01, DRB1 * 16:02, DRB1 * 03:01, DRB1 * 04:01, DRB1 * 04:04, DRB1 * 11:01, DRB1 * 12:01, DRB1 * 13:01, DRB1 * 13:02, DRB1 * 14:01, DRB1 * 14:02, DRB1 * 07:01, DRB1 * 08:01, DRB1 * 08:02, DRB1 * 08:03, DRB1 * 09:01, and DRB1 * Including but not limited to 10:01.

[0220] In some embodiments, the MHC class I molecule is HLA-A * 02. HLA-A * 01. HLA-A * 03. HLA-A * 11. HLA-A * 23. HLA-A * 24. HLA-B * 07. HLA-B * 08. HLA-B * 40, HLA-B * 44. HLA-B * 15. HLA-C * 04. HLA * C * 03, and HLA-C * 07. Allelic variants of the above HLA types also exist, all of which are encompassed by the present disclosure.

[0221] In some embodiments, the MHC molecule is HLA-A * 02:01 or HLA-A * It could be 01:01.

[0222] Naturally occurring MHC class I molecules consist of an α (heavy) chain associated with a β2-microglobulin. The heavy chain consists of subunits α1-α3. The β2-microglobulin protein and the α3 subunit of the heavy chain are associated. In certain embodiments, the β2-microglobulin and the α3 subunits are covalently linked. In certain embodiments, the β2-microglobulin and the α3 subunits are non-covalently linked. The α1 and α2 subunits of the heavy chain fold together to form a groove for peptides to be displayed and recognized by the TCR.

[0223] In some embodiments, the MHC comprised in the pMHC complexes of the present disclosure comprises (i) a class I MHC polypeptide, or a fragment, variant, or derivative thereof, and, optionally, (ii) a β2 microglobulin polypeptide, or a fragment, variant, or derivative thereof. In one specific embodiment, the class I MHC polypeptide is linked to the β2 microglobulin polypeptide by a peptide linker.

[0224] The pMHC complexes of the present disclosure may be in isolated and / or substantially pure form. For example, the complexes may be provided in a form that is substantially free of other peptides or proteins. The MHC molecules disclosed herein may include recombinant MHC molecules, non-naturally occurring MHC molecules, and functionally equivalent fragments of MHC, including derivatives or variants thereof, provided that the peptide bonds are maintained. For example, the MHC molecules may be attached to a solid support, in soluble form, attached to a tag, in biotinylated form, and / or in multimeric form. The peptides disclosed herein may be covalently attached to the MHC.

[0225] Methods for producing soluble recombinant MHC molecules to which the peptides disclosed herein can form complexes include, but are not limited to, expression and purification from E. coli cells or insect cells. Alternatively, MHC molecules can be produced synthetically or using cell-free systems.

[0226] The peptides disclosed herein may be presented in a complex with MHC on the cell surface. Thus, the present disclosure also provides a cell that presents the pMHC complexes disclosed herein on its surface. Such cells may be mammalian cells, preferably cells of the immune system, and professional antigen presenting cells (APCs), such as dendritic cells or B cells. Other preferred cells include T2 cells. The cells presenting the peptides or pMHC complexes of the present disclosure may be isolated, preferably in the form of a homogenous population, or may be provided in a substantially pure form. Such cells may not naturally present the pMHC complexes of the present disclosure, or the cells may present the pMHC complexes at a higher level than would be expected in nature. Such cells may be obtained by pulsing the aforementioned cells with one or more peptides of the present disclosure (e.g., 2-10, 2-20, 2-30, 5-25, 5-20, or 10-15 peptides) or genetically modifying the aforementioned cells (by transfection of DNA or RNA) to express one or more peptides of the present disclosure (e.g., 2-10, 2-20, 2-30, 5-25, 5-20, or 10-15 peptides). Pulsing is typically performed at 10 -5 From 10 -12 The method includes incubating the cells with the peptide for several hours, using peptide concentrations ranging from 0.1 to 100 M. The cells are then incubated with the peptide for several hours, using peptide concentrations ranging from 0.1 to 100 M. * The cells may be further transduced with a peptide-presenting marker (e.g., 02) to further induce peptide presentation. The cells may be recombinantly produced. The cells presenting the peptides of the present disclosure may be used to isolate antigen-binding molecules (e.g., antibodies, T cells, TCRs, and CARs) that can bind to the cells.

[0227] The peptides or pMHC complexes disclosed herein may be fused or conjugated to one or more heterologous molecules.The peptides or pMHC complexes disclosed herein may also be in the form of multimers.Thus, the present disclosure also provides fusion proteins, conjugates, and oligomeric complexes comprising the peptides or pMHC complexes disclosed herein.

[0228] In some embodiments, the peptide is fused or conjugated to one or more heterologous molecules, which may include an MHC molecule (or a fragment thereof).

[0229] Heterologous molecules suitable for genetic fusion and / or chemical conjugation to a peptide or pMHC complex of the present disclosure include, but are not limited to, peptides, polypeptides, small molecules, polymers, nucleic acids, lipids, sugars, etc. The heterologous molecule may be fused to the N-terminus and / or C-terminus of the peptide and / or another polypeptide chain in the pMHC complex.

[0230] Heterologous peptides and polypeptides include, but are not limited to, an epitope (e.g., FLAG) or tag sequence (e.g., His6, etc.) to allow for detection and / or isolation of the fusion protein; a transmembrane receptor protein or portion thereof (such as the extracellular domain or the transmembrane domain and the intracellular domain); a ligand or portion thereof that binds to a transmembrane receptor protein; a catalytically active enzyme or portion thereof; a polypeptide or peptide that promotes oligomerization (such as a leucine zipper domain); a polypeptide or peptide that increases stability (such as an immunoglobulin constant region (e.g., Fc domain)); a half-life extending sequence that includes a combination of two or more (e.g., 2, 5, 10, 15, 20, 25, etc.) naturally occurring or non-naturally occurring charged and / or uncharged amino acids (e.g., Ser, Gly, Glu, or Asp) designed to form a predominantly hydrophilic or predominantly hydrophobic fusion partner for the fusion protein; a functional or non-functional antibody (e.g., an antibody specific for dendritic cells), or a heavy or light chain thereof; and a polypeptide having an activity different from the fusion protein of the present disclosure.

[0231] In some embodiments, the fusion proteins of the present disclosure may include one or more affinity tags, for example, to allow affinity purification or coupling of another molecule. Examples of affinity tags include, but are not limited to, His6 tag, Avi-tag, biotin, hemagglutinin (HA) tag, FLAG tag, Myc tag, GST tag, MBP tag, chitin-binding protein tag, calmodulin tag, V5 tag, streptavidin-binding tag, green fluorescent protein (GFP), YFP, RFP, CFP, mCherry, tdTomato, SUMO tag, and ubiquitin tag.

[0232] The peptides or pMHC complexes of the present disclosure may be provided in soluble form or may be immobilized by attachment to a suitable solid support. Examples of solid supports include, but are not limited to, beads, membranes, sepharose, magnetic beads, plates, tubes, and columns. The pMHC complexes may be attached to ELISA plates, magnetic beads, or surface plasmon resonance biosensor chips. Methods for attaching peptides or pMHC complexes to solid supports are known to those skilled in the art and include, for example, the use of affinity binding pairs (e.g., biotin and streptavidin, or antibodies and antigens). In some embodiments, the peptides or pMHC complexes are labeled with biotin and attached to a streptavidin-coated surface.

[0233] In another aspect, the disclosure provides an isolated polynucleotide comprising a nucleic acid sequence encoding one or more of the peptides and / or peptide-based molecules of the disclosure, such as a complex (e.g., a pMHC complex), a fusion protein, or a conjugate comprising the described peptides. The polynucleotide can be, for example, DNA, cDNA, PNA, RNA, or a combination thereof (either single-stranded and / or double-stranded), or a polynucleotide in a native or stabilized form (e.g., a polynucleotide having a phosphorothioate backbone, etc.), and may or may not contain introns, so long as it encodes a peptide.

[0234] In some embodiments, the polynucleotides described herein encode a peptide comprising the amino acid sequence of any one of SEQ ID NOs: 1-7, 9, and 11-74, or a fragment or derivative thereof. In some embodiments, the polynucleotides described herein encode a peptide comprising the amino acid sequence of any one of SEQ ID NOs: 1-7, 9, and 11-74, or a fragment or derivative thereof.

[0235] In a further aspect, the present disclosure provides a vector comprising the nucleic acid sequence of the present disclosure. The vector may comprise one or more additional nucleic acid sequences encoding one or more additional peptides in addition to the nucleic acid sequence encoding only the peptide of the present disclosure. Such additional peptides may be fused to the N-terminus or C-terminus of the peptide of the present disclosure when expressed. Examples of such additional peptides are detailed in the above section. In one embodiment, the vector comprises a nucleic acid sequence encoding a peptide or protein tag (such as a biotinylation site, a FLAG-tag, a MYC-tag, an HA-tag, a GST-tag, a Strep-tag, or a poly-histidine tag).

[0236] The off-target peptides identified herein can be used to evaluate and / or screen therapeutic molecules for molecules with minimal off-target effects. Such therapeutic molecules can include antigen recognition molecules (such as T cell receptors (TCRs), chimeric antigen receptors (CARs), antibodies, or antigen-binding fragments thereof). In various embodiments of the present disclosure, the antigen recognition molecules are present in solution. In various embodiments of the present disclosure, the antigen recognition molecules are present on cells (e.g., T cells, B cells, or hybridomas).

[0237] In one aspect, the disclosure provides an in vitro method of assessing off-target effects of an antigen recognition molecule, the method comprising: (a) contacting the antigen recognition molecule with a target peptide presented in a complex with a major histocompatibility complex (MHC) molecule (MHC-target peptide complex); (b) contacting the antigen recognition molecule with one or more off-target peptides associated with the target peptide, each of the off-target peptides being presented in a complex with the same MHC molecule as in (a) (MHC-off-target peptide complex); and (c) determining and comparing the binding level of the antigen recognition molecule to the MHC-target peptide complex and each of the MHC-off-target peptide complexes.

[0238] In one aspect, the disclosure provides an in vitro method of assessing an off-target effect of an antigen recognition molecule, comprising: a) contacting said antigen recognition molecule with one or more off-target peptides related to a target peptide recognized by said antigen recognition molecule, wherein each of said off-target peptides is presented in a complex with a major histocompatibility complex (MHC) molecule (MHC-off-target peptide complex); and b) determining the level of binding of said antigen recognition molecule to each of said MHC-off-target peptide complexes.

[0239] In some embodiments of the in vitro method described above, the binding level is determined by detecting the amount of binding of the antigen recognition molecule to the MHC-peptide complex. For example, if the signal to noise ratio (such as that shown in Table 7 in Example 2) of the signal that correlates with the binding of the antigen recognition molecule to the MHC-peptide complex is at least about 2.0, 2.5, 3.0, 3.5, 4.0, or 5.0, the antigen recognition molecule may be determined to exhibit detectable binding to the MHC-peptide complex. If the antigen recognition molecule specifically binds to one or more MHC-off-target peptide complexes, the antigen recognition molecule may exhibit off-target effects. In some specific embodiments, an signal to noise ratio of at least about 3 indicates detectable binding. In some embodiments, the binding affinity may be determined. The binding affinity may be determined by measuring the equilibrium dissociation constant (KD) of the binding reaction. Alternatively, the binding affinity may be characterized using other methods.

[0240] In some embodiments, the in vitro method comprises determining that an antigen recognition molecule is likely to have an off-target effect if the antigen recognition molecule detectably binds to at least one MHC-off-target peptide complex, wherein the off-target peptide is expressed in an essential normal tissue.

[0241] In another aspect, the present disclosure provides a method for selecting an antigen recognition molecule that binds to a target pMHC complex with minimal off-target effects. Such a method can include the steps of: (a) contacting a plurality of antigen recognition molecules with a target peptide presented in a complex with a major histocompatibility complex (MHC) molecule (MHC-target peptide complex); (b) contacting the same plurality of antigen recognition molecules with one or more off-target peptides associated with the target peptide, each of the off-target peptides being presented in a complex with the same MHC molecule as in (a) (MHC-off-target peptide complex); c) selecting one or more antigen recognition molecules based at least in part on the number of MHC-off-target peptide complexes detectably bound by each of the antigen recognition molecules; and d) optionally repeating steps (a)-(c) using the selected antigen recognition molecules.

[0242] In some embodiments, the method includes selecting an antigen recognition molecule that binds to a minimal number of MHC-off-target peptide complexes. In some embodiments, the selected antigen recognition molecule detectably binds to no more than 5 (e.g., no more than 4, no more than 3, no more than 2, or no more than 1) MHC-off-target peptide complexes, where the off-target peptides are expressed in essential normal tissues.

[0243] In some embodiments, the selected antigen recognition molecule does not detectably bind to any MHC-off-target peptide complexes, where said off-target peptides are expressed in essential normal tissues.

[0244] In some embodiments, when some selected antigen recognition molecules bind to at least one MHC-off-target peptide complex, the method may also include comparing the binding level of the antigen recognition molecule to the MHC-target peptide complex and the binding level to the MHC-off-target peptide complex. The antigen recognition molecule may be more likely to bind to the MHC-target peptide complex than to the MHC-off-target peptide complex. For example, the antigen recognition molecule may bind to the MHC-target peptide complex about 1000 times, about 500 times, about 200 times, about 100 times, about 90 times, about 80 times, about 70 times, about 60 times, about 50 times, about 40 times, about 30 times, about 20 times, or about 10 times more strongly to the MHC-target peptide complex than to the MHC-off-target peptide complex. The method may further include selecting an antigen recognition molecule based at least in part on an MHC-target peptide complex / MHC-off-target peptide complex binding ratio for one or more off-target peptides, where a higher MHC-target peptide complex / MHC-off-target peptide complex binding ratio is more desirable.

[0245] In some embodiments, the selection of an antigen recognition molecule may take into account the binding level of multiple antigen recognition molecules to the MHC-target peptide complex, where stronger binding to the MHC-target peptide complex is more desirable. Thus, the method may include selecting an antigen recognition molecule based at least in part on the binding level to the MHC-target peptide complex.

[0246] In some embodiments, the selected antigen recognition molecule binds to at least one potential secondary target peptide in addition to the MHC-target peptide complex.

[0247] Methods for determining binding to pMHC complexes include, for example, surface plasmon resonance (e.g., BIACORE™), or any other biosensor technology, enzyme-linked immunosorbent assay (ELISA), ELISpot, luminescence assay, flow cytometry, chromatography, or microscopy. Alternatively, or in addition, binding may be determined by a functional assay in which a biological response (e.g., cytokine release or cell apoptosis) is detected upon binding.

[0248] For example, antibodies and TCRs may be obtained from display libraries where the library is panned using the pMHC complexes of the present disclosure. TCRs may be displayed on the surface of phage particles and yeast particles, and such libraries have been used, for example, to isolate high affinity variants of TCRs derived from T cell clones. TCR phage libraries may be used to isolate TCRs with novel antigen specificities. Such libraries may be constructed using α- and β-chain sequences that correspond to those found in natural repertoires. However, the random combination of these α- and β-chain sequences that occurs during library creation may produce a repertoire of TCRs that may not exist in nature.

[0249] In some embodiments, the pMHC complexes of the present disclosure may be used to screen a library of diverse TCRs displayed on the surface of phage particles. The TCRs displayed by said libraries may not correspond to those contained in the natural repertoire, for example, the TCRs may include α- and β-chain pairings that are not believed to exist in vivo, and / or the TCRs may include non-natural mutations, and / or the TCRs may be in a soluble form. The screening may involve panning the phage library with the pMHC complexes of the present disclosure and subsequent isolation of the bound phage particles. For this purpose, the pMHC complexes may be attached to a solid support (such as magnetic beads, or a column matrix), and the isolated phage-bound pMHC complexes may be isolated using a magnet or by chromatography, respectively. The panning process may be repeated several times. The isolated phage may be further expanded in E. coli cells. The isolated phage particles may be tested for specific binding to the pMHC complexes of the present disclosure. Binding can be detected using techniques described herein, such as, but not limited to, ELISA or SPR, e.g., using a BIACORE™ instrument. The DNA sequence of the T cell receptor displayed by the pMHC-binding phage can be further identified by PCR methods.

[0250] Alternatively, antigen-binding T cells and TCRs can be isolated from fresh blood obtained from patients or healthy donors. Such methods include stimulating T cells using autologous dendritic cells (DCs) followed by autologous B cells, and then pulsing with the target / off-target peptides disclosed herein. Several rounds of stimulation (e.g., 3 or 4 rounds) may be performed. The activated T cells may then be tested for recognition of the target / off-target peptides by measuring cytokine release in the presence of T2 cells pulsed with the target / off-target peptides of the present disclosure (e.g., using an IFNγ ELISpot assay). The activated cells may then be sorted by fluorescence activated cell sorting (FACS) using labeled antibodies to detect intracellular cytokine production (e.g., IFNγ) or expression of cell surface markers (such as CD137). The sorted cells may be expanded and further confirmed, for example, by ELISpot assay and / or cytotoxicity against target cells and / or staining with peptide-MHC tetramers. The TCR chains from the validated clones can then be amplified by rapid amplification of cDNA ends (RACE) and sequenced. An exemplary method for isolating and validating TCRs specific to a target antigen is described in Moore et al., Sci. Immunol. 6, eabj4026 (2021) (incorporated herein by reference in its entirety).

[0251] TCR screening can also be performed using a TCR activation assay. For example, JRT3-T3.5 cells (ATCC TIB-153), a Jurkat subline lacking endogenous TCR surface expression, can be utilized as described in Moore et al., Sci. Immunol. 6, eabj4026 (2021), which is incorporated by reference in its entirety. T cell receptor alpha (TCRA) and T cell receptor beta (TCRB) sequences of interest can be introduced into cells by lentiviral transduction, and surface TCR+ cells can be selected. Antigen-presenting cells (e.g., 293T cells) pulsed with target or off-target peptides can be incubated with TCR-transduced or parental JRT3 cells. Readouts (e.g., luciferase activity) are then measured as an indication of TCR-mediated activation.

[0252] In some embodiments, monoclonal antibody screening can be performed using cells isolated from spleen and lymphoid tissues harvested from mice with optimal titers using hybridoma and B cell sorting (BST) platforms. A counter-screening approach using one or more off-target peptides can help identify and eliminate B cells and hybridomas that show cross-reactivity with peptides that form pHLA complexes similar to the targeted complex. For example, antigen positive (Ag) B cells with cross-reactivity to off-target peptides can be used to screen for B cells and hybridomas that show cross-reactivity with off-target peptides. + ) clones can be identified by testing cell supernatants for antibody binding to cells (e.g., T2 cells) pulsed with the target peptide or off-target peptide using a cell binding assay. + B cells can be captured using a biotinylated HLA-target peptide complex in the presence of a high concentration of one or more unlabeled HLA-off-target peptide complexes to enrich for antibodies specific to the HLA-target peptide complex. +The antibody variable domains from the B cells can then be cloned as full-length mAbs and expressed (eg, in CHO cells) for further screening.

[0253] Antibodies that bind to target / off-target peptides can be determined using ELISA. For example, MHC-target peptide complexes or MHC-off-target peptide complexes can be coated on a plate (e.g., a 96-well microtiter plate). A sample containing test antibodies can be added to the plate, and the reaction can be incubated under binding conditions. The plate can then be washed, and a secondary antibody can then be added to the plate to detect the antibody that binds to the MHC-peptide complex. Typically, the secondary antibody can generate a signal that indicates the amount of antibody binding to the MHC-peptide complex.

[0254] In various embodiments of the methods described herein, the pMHC complexes may be provided in soluble form or may be immobilized by attachment to a suitable solid support. Examples of solid supports include, but are not limited to, beads, membranes, sepharose, magnetic beads, plates, tubes, columns. The pMHC complexes may be attached to an ELISA plate, magnetic beads, or a surface plasmon resonance biosensor chip. Methods for attaching the pMHC complexes to a solid support are known to those skilled in the art and include, for example, the use of affinity binding pairs (e.g., biotin and streptavidin, or antibodies and antigens). In some embodiments, the pMHC complexes are labeled with biotin and attached to a streptavidin-coated surface.

[0255] In various embodiments of the methods described herein, the pMHC complex may be present on a cell. Such cells may be mammalian cells, preferably cells of the immune system, and professional antigen presenting cells (APCs), such as dendritic cells or B cells. Other preferred cells include T2 cells. Cells presenting peptides or pMHC complexes of the present disclosure may be isolated, preferably in the form of a homogenous population, or provided in substantially pure form. Such cells may be obtained by pulsing said cells with one or more peptides of the present disclosure (e.g., 2-10, 2-20, 2-30, 5-25, 5-20, or 10-15 peptides) or by genetically modifying said cells (by transfer of DNA or RNA) to express one or more peptides of the present disclosure (e.g., 2-10, 2-20, 2-30, 5-25, 5-20, or 10-15 peptides). Pulsing is typically performed over a period of 10 minutes. -5 From 10 -12 The method includes incubating the cells with the peptide for several hours, using peptide concentrations ranging from 0.1 to 10 M. The cells are then incubated with the peptide for several hours, using peptide concentrations ranging from 0.1 to 10 M. * The cells may be further transduced with a recombinant vector (e.g., 02) to further induce presentation of the peptide. The cells may be recombinantly produced.

[0256] In various embodiments of the methods described herein, the methods are performed in a high-throughput format (eg, 96-well plates).

[0257] In yet another aspect, provided herein is a method for enriching a sample with antigen recognition molecules that specifically bind to a target peptide, comprising: (a) contacting a sample containing a plurality of antigen recognition molecules with the target peptide in the presence of one or more off-target peptides associated with the target peptide, wherein each of the target peptide and the one or more off-target peptides is presented in a complex with a major histocompatibility complex (MHC) molecule (MHC-target peptide complex or MHC-off-target peptide complex); and (b) enriching the sample by isolating the antigen recognition molecules bound to the MHC-target peptide complex. The method may further comprise repeating steps (a)-(b) to further enrich the sample.

[0258] There are various methods that can isolate the antigen recognition molecule bound to the MHC-target peptide complex. For example, the MHC-target peptide complex can be present on the antigen presenting cell, while the MHC-off-target peptide complex is not present on the antigen presenting cell (e.g., in soluble form). Alternatively, the MHC-target peptide complex can be immobilized on a solid support, while the MHC-off-target peptide complex can be soluble or immobilized on a different solid support. Also, the MHC-target peptide complex and the MHC-off-target peptide complex can be differentially labeled so that the MHC-target peptide complex and the antigen recognition molecule bound to the aforementioned complex can be specifically detected and isolated. The antigen recognition molecule can be eluted from the MHC-target peptide complex after isolation.

[0259] Certain embodiments and implementations of the disclosed technology are described above with reference to block diagrams and flow diagrams of systems and methods and / or computer program products according to exemplary embodiments or implementations of the disclosed technology. It will be understood that one or more blocks of the block diagrams and flow diagrams, and combinations of blocks in the block diagrams and flow diagrams, respectively, can be implemented by computer-executable program instructions. Similarly, some blocks of the block diagrams and flow diagrams may not necessarily be performed in the order shown, may be repeated, or may not necessarily be performed at all according to some embodiments or implementations of the disclosed technology.

[0260] These computer-executable program instructions may be loaded onto a general purpose computer, special purpose computer, processor, or other programmable data processing apparatus to produce a particular machine, such that the instructions executing on the computer, processor, or other programmable data processing apparatus create means for implementing one or more functions specified in the flow chart blocks. Also, these computer program instructions may be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture that includes instruction means for implementing one or more functions specified in the flow chart blocks.

[0261] As an example, an embodiment or implementation of the disclosed technology may provide a computer program product including a computer usable medium having computer readable program code or program instructions embodied thereon, the computer readable program code being adapted to perform implementation of one or more functions specified in the flow chart blocks. Similarly, the computer program instructions may be loaded into a computer or other programmable data processing apparatus such that a sequence of operational elements or steps are executed on the computer or other programmable apparatus to execute a process on the computer, such that the instructions executing on the computer or other programmable apparatus provide elements or steps for implementing the functions specified in the flow chart blocks.

[0262] Thus, the blocks of the block diagrams and flow diagrams represent combinations of means for performing specified functions, combinations of elements or steps for performing specified functions, and program instruction means for performing specified functions. It will also be understood that each block of the block diagrams and flow diagrams, and combinations of blocks in the block diagrams and flow diagrams, can be implemented by special purpose hardware-based computer systems that perform the specified functions, elements, or steps, or combinations of special purpose hardware and computer instructions.

[0263] Certain implementations of the disclosed technology are described above with reference to customer devices that may include handheld computing devices. Those skilled in the art will recognize that there are several categories of handheld devices, which are commonly known as portable computing devices and may be battery powered, but are not typically classified as laptops. For example, handheld devices may include, but are not limited to, portable computers, tablet PCs, Internet tablets, PDAs, ultra-mobile PCs (UMPCs), wearable devices, and smartphones. Additionally, implementations of the disclosed technology may be utilized with Internet of Things (IoT) devices, smart televisions and media devices, home appliances, automobiles, toys, and peripherals that interface with these devices using voice command devices.

[0264] While certain embodiments of the present disclosure have been described in conjunction with what are presently considered to be the most practical and various embodiments, it is to be understood that the disclosure is not limited to the disclosed embodiments, but rather is intended to cover various modifications and equivalent arrangements included within the scope of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and are not intended to be limiting of the invention.

[0265] This specification uses examples to disclose certain embodiments of the technology and to enable any person skilled in the art to practice certain embodiments of the technology, including making and using any devices or systems and practicing any incorporated methods. The patentable scope of certain embodiments of the technology is defined in the claims, and may include other examples that occur to those skilled in the art. Such other examples are intended to be within the scope of the claims if they have structural elements that do not differ from the literal language of the claims, or if they include equivalent structural elements that do not differ substantially from the literal language of the claims. EXAMPLES

[0266] Working Example The following examples are provided to further illustrate some of the embodiments disclosed herein. The examples are intended to illustrate, but not limit, the disclosed embodiments. Example 1. Development of Peptide in Groove Similarity Prediction (PIGSPRED) to predict off-targets associated with target peptide-HLA (PHLA) complexes

[0267] Peptide-HLA (pHLA) complexes that are specifically expressed on cancer cells present a unique set of targets for destroying cancer cells using antibody-based or cell-based therapeutic approaches. However, it is important to consider potential off-targets associated with these pHLA complexes to avoid off-target toxicity. Although the expression of these targets should be specific to cancer cells, it is becoming clear that another set of off-targets exists: pHLA complexes that are highly similar to the target pHLA (Cameron BJ et al., Sci Transl Med. 2013; Linette GP et al., Blood. 2013). The three-dimensional nature of the interaction between pHLA and its cognate therapeutic molecule makes it difficult to elucidate the similarity of pHLA complexes. However, most reported off-target toxicity was due to off-targets that share the same HLA as the target pHLA and also contain peptides similar / homologous to peptide sequences in the target pHLA. This example describes the development of a method called PIGSPRED (Peptide in Groove Similarity Prediction), which is useful in predicting such off-targets. After identifying the target pHLA of interest, PIGSPRED predicted the off-targets associated with the target, which was useful in assessing the risk associated with the target during the target selection step and in screening for highly specific T cell receptors (TCRs) or antibodies (Abs). After inputting the target pHLA, the PIGSPRED method operated in the following multi-step manner. 1. Identifying peptide locations important for TCR (T cell receptor) or antibody binding

[0268] When assessing the similarity / homology of a peptide to a target peptide, it is necessary to evaluate the similarity at positions of the peptide that may be involved in binding interactions with the TCR / antibody. These important positions can be confirmed by analyzing experimental structures of pHLA in complex with the TCR / antibody, typically derived using crystallography or cryo-electron microscopy (cryoEM) techniques. Not only are such structures difficult to obtain, but at the initial target selection stage, when assessing potential risks associated with the target, the TCR / antibody for the target is not available. To address this challenge, we implemented a computational mutagenesis algorithm using a commercially available tool called NetMHCpan, which can predict the binding affinity of peptides to HLA using machine learning models (Jurtz V et al., J Immunol. 2017). First, the binding affinity of the peptide to HLA was predicted. Then, the amino acid at each position in the peptide was iteratively mutated to a glycine amino acid, and the affinity of the mutated peptide was predicted thereafter. Positions that did not lose binding affinity to HLA were flagged for further consideration, as they were unlikely to be involved in HLA binding and therefore freely interacted with the TCR / antibody. Among these free positions, those containing non-glycine amino acids were identified as important positions for TCR / antibody binding. As shown above, when the structure of pHLA in complex with TCR / antibody was available, binding motifs in the peptide sequence involved in the TCR / antibody interaction could be obtained. 2. Identification of Similar Peptides

[0269] Standard human protein sequences were obtained from UniprotKB to identify similar peptides that were the same length as the target peptide and were identical to the target peptide at three or more critical amino acid positions. Similar peptides were further identified based on experimentally derived binding motifs in the target peptide sequence. If these similar peptides were potential off-targets, the aforementioned peptides would have to be able to bind to HLA in the target pHLA. Therefore, NetMHCpan was used to predict the binding affinity of each similar peptide to HLA, and peptides that were not predicted to bind were discarded from the database. Similar peptides may be identical to the target peptide, but from different protein sources than the target peptide. 3. Expression Analysis

[0270] Upon identification of similar peptides, it was necessary to establish that the aforementioned peptides could potentially produce off-target effects in order to avoid overestimation of the risk associated with a given target. Therefore, we confirmed the expression of the peptides in essential normal tissues or essential cell types. For this purpose, we queried the Gene Tissue Expression Database or GTEx (gtexportal.org / home / ). The GTEx database version 6 is based on genome build GRCh38 and contains gene expression data across 51 tissue types from 549 healthy donors. We queried gene expression values ​​only in tissue types considered essential (in other words, tissues that cannot be sacrificed to save the patient's life, unlike non-essential tissues (e.g., breast, ovary, testis, etc.)). For each gene, we calculated the 95th percentile expression value in each essential tissue type, followed by the maximum of the aforementioned values ​​across all essential tissue types. We assumed that any similar peptide derived from the aforementioned gene was a potential off-target if the maximum expression value of the gene exceeded 0.5 transcripts per millicon (TPM). 4. Final output

[0271] Once all potential off-targets have been identified, a degree of similarity (DoS) score for each off-target peptide was calculated to quantify the homology between the off-target and target peptides. DoS represents the number of identical amino acids at the same position or Hamming distance between peptides. Once the DoS score was determined, the number of off-targets at different thresholds was calculated to compare the likelihood of off-target toxicity associated with different targets. The immunopeptidemics data generated in this example was used to annotate the off-target peptides observed in mass spectrometry experiments, thereby providing further evidence for potential off-targets. Identification of potential off-targets associated with HLA-A2 peptides Identification of potential off-targets associated with the first MAGEA4 targeting peptide: GVYDGREHTV (SEQ ID NO: 1)

[0272] Tables 1A to 3C show MAGEA4 in step 102. 230~239 Target GVYDGREHTV (SEQ ID NO: 1)-HLA-A * The results of applying the PIGSPRED method to input 02:01 are shown below.

[0273] Table 1A shows the MAGEA4 performance at different DoS thresholds. 230~239 Target GVYDGREHTV (SEQ ID NO: 1)-HLA-A * 02:01 includes the number of potential high-risk off-target peptides associated with 02:01. The number of potential high-risk off-target peptides identified by the computer is provided. The critical positions were predicted by the computer in step 104 to be positions 4 and 6-9. The table also provides the number of potential high-risk off-target peptides predicted by the computer that were also identified in the immunopeptidomics mass spectrometry data. Table 1A: [Table 1A]

[0274] Table 1B shows the MAGEA4 DoS of 6 or more230~239 Target GVYDGREHTV (SEQ ID NO: 1)-HLA-A * The predicted binding affinity of each off-target was calculated using the half maximal inhibitory concentration (IC 50 ) value and binding affinity percentile rank value (%Rank_BA in NetMHCpan). The mRNA level in the highest expressed normal tissue for each off-target gene (TMP from GTEx) is provided, and the top three highly expressed normal tissues are also provided. The peptides found in mice may be useful for performing antigen recognition molecule binding studies (e.g., screening) in mice. Potential off-target (1) GLADGRTHTV (SEQ ID NO:2) was detected in the mass spectrometry experiment, while other potential off-targets in the table were not detected. Potential off-target peptide GLYDGREHSV (SEQ ID NO:73, DoS=8) from MAGEA8 gene (TPM=0.4, not expressed in mouse) and potential off-target peptide GLYDGMEHLI (SEQ ID NO:74, DoS=6) from MAGEA10 gene (TPM=0.1, not expressed in mouse) were filtered from the list of potential off-target peptides in Table 1B due to their low expression levels in essential normal tissues. However, since both MAGEA8 and MAGEA10 are cancer / testis (CT) antigens that may be useful for targeting in cancer immunotherapy applications, a separate working list may be prepared containing such potential secondary targets, and further testing (e.g., antigen recognition molecule screening) may be performed on one or more such potential secondary targets as necessary to assess whether cross-reactivity with the aforementioned peptides may provide a beneficial therapeutic effect. Table 1B: [Table 1B-1] [Table 1B-2]

[0275] Table 2A shows the MAGEA4 performance at different DoS thresholds. 230~239 Target GVYDGREHTV (SEQ ID NO: 1)-HLA-A * Includes a number of potential off-targets associated with 02:01. An empirically derived binding motif was used: XXXDXREXXX (SEQ ID NO: 8), where X represents any amino acid. Table 2A: [Table 2A]

[0276] Table 2B shows the MAGEA4 with DoS>=3. 230~239 Target GVYDGREHTV (SEQ ID NO: 1)-HLA-A * Potential off-targets related to 02:01 are included. An experimentally derived binding motif XXXDXREXXX (SEQ ID NO: 8) was used, where X represents any amino acid. Potential off-target (9) ALVDQRELYL (SEQ ID NO: 9) was detected in the mass spectrometry experiment, while other potential off-targets in the table were not detected. Table 2B: [Table 2B]

[0277] Table 3A shows the MAGEA4 performance at different DoS thresholds. 230~239 Target GVYDGREHTV (SEQ ID NO: 1)-HLA-A * The number of potential off-targets associated with 02:01 is included. An empirically derived binding motif was used: XX[YFWM][DENQ][GA][R][DENQ]XXX (SEQ ID NO: 10), where X represents any amino acid. Table 3A: [Table 3A]

[0278] Table 3B shows the MAGEA4 with DoS>=2. 230~239 Target GVYDGREHTV (SEQ ID NO: 1)-HLA-A* Potential off-targets related to 02:01 were included. An experimentally derived binding motif was used, XX[YFWM][DENQ][GA][R][DENQ]XXX (SEQ ID NO: 10), where X represents any amino acid. Potential off-target (5)ALVDQRELYL (SEQ ID NO: 9) was detected in the mass spectrometry experiment, while other potential off-targets in the table were not detected. Table 3B [Table 3B-1] [Table 3B-2] Identification of potential off-targets associated with the following second MAGEA4 targeting peptide: KVLEHVVRV (SEQ ID NO: 48)

[0279] Tables 4A-4B show alternative MAGEA4 peptide-derived pMHC targets in step 102, MAGEA4 286~294 Target KVLEHVVRV (SEQ ID NO: 48)-HLA-A * The results of applying the PIGSPRED method when 02:01 is input are shown below.

[0280] Table 4A shows the MAGEA4 performance at different DoS thresholds. 286~294 Target KVLEHVVRV (SEQ ID NO: 48)-HLA-A * The table includes the number of potential high risk off-target peptides associated with 02:01. The number of potential high risk off-target peptides identified by computation is provided. The critical positions were computationally predicted in step 104. The table also provides the number of potential high risk off-target peptides predicted by computation that were also identified in the immunopeptidomics mass spectrometry data. Table 4A: [Table 4A]

[0281] Table 4B includes potential off-targets associated with targets with a DoS of 6 or more. Eleven different potential off-targets were detected in the mass spectrometry experiments, while the other potential off-targets in the table were not detected. 286~294 The off-target peptide KVLEHVVRV (sequence number 49, DoS=9) from the MAGEA8 gene, which is identical to the target peptide (TPM=0.4, not expressed in mice), was filtered from the list of potential off-target peptides in Table 4B due to its low expression levels in essential normal tissues. Table 4B: [Table 4B-1] [Table 4B-2]

[0282] MAGE-A4 230~239 and MAGE-A4 286~294 Both were computationally predicted to bind HLA-A02 with relatively high affinity (IC 50 = 560.08 nM and 8.52 nM). The PIGSPRED method shown in Tables 1B and 4B was applied to MAGE-A4 230~239 and MAGE-A4 286~294 Six and 23 peptides with DoS >6 that are expressed in normal essential tissues were identified for MAGE-A4, suggesting that targeting the former may be less likely to result in off-target toxicity. 230~239 and MAGE-A4 286~294 Both share a high DoS (DoS 9 and 8, respectively) against a peptide derived from MAGE-A8, another CT antigen that shows negligible expression in essential normal tissues. As above, with respect to Table 1B, such off-targets can be tabulated separately and, if necessary, tested for binding of target antigen recognition molecules. Identification of potential off-targets associated with HLA-A1 peptides

[0283] Using the in-silico computational strategy detailed above, in step 102, HLA-A * MAGEA3, a target peptide related to 01:01 168~176 Target EVDPIGHLY (SEQ ID NO: 29)-HLA-A * Potential off-targets associated with 01:01 were identified.

[0284] Table 5A shows the MAGEA3 performance at different DoS thresholds. 168~176 Target EVDPIGHLY (SEQ ID NO: 29)-HLA-A * The number of potential off-targets associated with 01:01 is shown. Important positions were computationally predicted in step 104. Table 5A: [Table 5A]

[0285] Table 5B shows the results for MAGEA3 with DoS>=5. 168~176 Target EVDPIGHLY (SEQ ID NO: 29)-HLA-A * A representative list of potential off-targets associated with 01:01 is shown. Additional off-targets with DoS >= 5 or less are not shown (indicated with "..."). Potential off-targets (1), (2), (3), (10), (12), (14)-(16), and (18) were detected in mass spectrometry experiments, while other potential off-targets in the table were not detected. Table 5B: [Table 5B-1] [Table 5B-2] [Table 5B-3]

[0286] Of note, the peptide ESDPIVAQY (SEQ ID NO: 47) derived from the muscle protein titin inhibits MAGEA3 168~176 It was identified because it is one of the highly ranked potential off-target peptides for the target (see peptide 18 in Table 5B). Expression of this peptide is also found in cardiac tissue. Interestingly, this peptide has been reported in Cameron BJ et al., Sci Transl Med. 2013 as a cross-reactive target for engineered MAGE A3-directed T cells, and is most likely responsible for the in vivo cardiotoxicity observed in clinical trials evaluating engineered MAGE A3-directed T cells. This result demonstrates that the prediction method described herein can accurately identify potential off-targets for pHLA complexes and can be applied to reduce the risk of off-target toxicity in future clinical investigations. Example 2. MAGEA4 230~239 and anti-HLA-A2:MAGEA4 to T2 cells pulsed with related off-target peptides 230~239 Antibody binding

[0287] Ab A and Ab B are human leukocyte antigen (HLA) class I alleles HLA-A * It is a monoclonal antibody (mAb) with a human fragment crystallizable (Fc) region that recognizes amino acids 230-239 of MAGEA4 when complexed with 02:01.

[0288] HLA-A * 02:01 Anti-HLA-A2:MAGEA4 to positive T2 (174 CEM.T2) cells 230~239 Antibody cell surface binding was assessed by a flow cytometry-based peptide pulsing assay. For pulsing, 1 × 10 6 T2 cells were cultured in 10 μg / ml human (h)B2M (EMD Millipore Cat. No. 475828) and 100 μg / ml MAGEA4 230~239The peptides were incubated in 1 ml of AIM V medium (Gibco. Cat. No. 31035025) for 16 hours at 37°C. Cells were washed with staining buffer (calcium and magnesium free PBS (Corning, ref. no. 21-031-CV) + 2% FBS (Seradigm, lot no. 238B15)), harvested using cell dissociation buffer (Millipore, Cat. No. S-004-C), and resuspended in staining buffer. Pulsed cells (200,000 cells) were plated in 96-well V-bottom plates (Axygen, Cat. No. P-96-450V-CS) and stained with 3-fold serial dilutions (1.7 pM to 100 nM) of Ab A, Ab B, non-binding isotype control antibody, or HLA-A2 antibody (data not shown in Table 6) for 30 minutes at 4°C. The cells were then washed once with staining buffer and incubated with 5 μg / ml Alexa Fluor 647 conjugated to Fab'2 anti-mouse Fc specific secondary antibody (Jackson ImmunoResearch, Cat. No. 115-606-071) for 30 min at 4°C. Finally, the cells were stained with green fluorescent viability dye (Molecular Probes Cat. No. L-34970, reconstituted in 50 μl DMSO) at a concentration of 1:1000. The cells were then washed and fixed using a 50% solution of BD Cytofix (BD, Cat. No. 554655) diluted in PBS. The samples were run on an intellicyt iQue flow cytometer (Intellicyt) and the results were analyzed using Forecyte analysis software (Intellicyte) to calculate mean fluorescence intensity (MFI) after gating on live cells. The MFI values ​​were plotted over a 12-point response curve in Graphpad Prism using a 4-parameter logistic equation to determine EC 50 Values ​​were calculated. Secondary antibody alone (i.e., no primary antibody) for each dose-response curve was also included in the analysis as a series of 3-fold dilutions and is expressed as the lowest dose. Signal to noise (S / N) was determined by taking the ratio of the highest MFI on the dose-response curve to the MFI in the secondary antibody only wells. EC 50 The values ​​(M) and maximum S / N are shown in Table 6. Ab A had an EC50 and Ab B bound with a maximum S / N of 365.9 and an EC of 1.3 nM. 50 and bound with a maximum S / N of 511.7. The isotype control antibody bound minimally with an S / N of 13.4. Table 6: MAGEA4 by flow cytometry 230~239 Anti-HLA-A2:MAGEA4 to T2 cells pulsed with 230~239 Antibody binding [Table 6] ND = EC because binding did not reach saturation within the antibody concentration range tested. 50 The value could not be determined accurately.

[0289] The in-silico computational strategy described in Example 1 above was used to determine the HLA-A * We identified several MAGEA4-associated peptides that are predicted to form complexes with 02:01. The identified peptides are summarized in Table 1B. Two HLA-A2:MAGEA4 binding domains for these associated peptides were identified. 230~239 Binding of the antibodies (Ab A and Ab B), a non-binding isotype control antibody, and the HLA-A2 antibody was assessed by the T2 pulsing assay described above. S / N is expressed as the ratio of pulsed cells to unpulsed (no peptide) cells. Peptide loading was determined when the signal-to-noise ratio for HLA-A2 binding was greater than 1. As summarized in Table 7, both Ab A and Ab B bound to the MAGEA4 peptide with S / N values ​​of 512.9 for Ab A and 747.2 for Ab B. Ab A bound to the MAGEA4 peptide with S / N values ​​of 512.9 for Ab A and 747.2 for Ab B. 230~239 While highly specific for the peptide, Ab B bound strongly to the off-target peptide LPIN2 with a S / N ratio of 27.5. There was no detectable binding to the remaining peptides, and control antibody binding was <4.3 for all peptides tested. Table 7: Anti-HLA-A2:MAGEA4 to T2 cells pulsed with relevant peptides via flow cytometry 230~239 Antibody binding [Table 7] Example 3. Application of PIGSPRED for target discovery and prioritization

[0290] Cancer-specific pHLA complexes can be identified by the confirmation of genes that are specifically expressed in cancer tissues. For this purpose, public databases including, for example, The Cancer Genome Atlas (TCGA) and the Genome Tissue Expression Database (GTEx) were used. Genes that were expressed in cancer types with a 75th percentile transcripts per millicon (TPM) value of more than 2 according to GTEx and were negligibly expressed in all essential normal tissues or essential cell types were classified as cancer-specific genes. Canonical protein sequences corresponding to cancer-specific genes were derived from the UniProtKB database and used to predict potential 8-12-mer peptide sequences predicted to bind to the HLA of interest. Predictions were made using the NetMHCpan tool. Once cancer-specific pHLAs were identified, PIGSPRED was used to calculate the number of potential off-targets associated with each cancer-specific pHLA. The number of potential off-targets is representative of the likelihood of off-target toxicity associated with the target, and therefore, this number was used to rank the list of pHLA targets and prioritize targets for the development of therapeutics. After target selection and generation of therapeutic molecules that bind to the target, the off-targets predicted by PIGSPRED play an important role in experimental screening of therapeutic molecules that do not bind to the off-targets, and thus the most specific therapeutic molecules are selected for further development. References [ka] [ka]

[0291] The present invention is not intended to be limited in scope by the specific embodiments described herein. Indeed, various modifications of the invention in addition to the embodiments described herein will become apparent to those skilled in the art from the foregoing description. Such modifications are intended to be within the scope of the appended claims.

[0292] All patents, applications, publications, test methods, articles, and other materials cited herein are incorporated by reference in their entirety as if physically illustrated herein.

Claims

1. A non-transitory computer-readable medium configured to communicate with one or more processors of a computing device, the non-transitory computer-readable medium, when executed by the processors, causing the computing device to: a) receiving as input a computational representation of a target peptide presented in complex with a major histocompatibility complex (MHC) molecule (an MHC-target peptide complex); b) predicting all amino acid positions within the target peptide that are available to interact with an antigen recognition molecule that recognizes the MHC-target peptide complex; d) generating a working list of peptides such that, within the total pool of predicted or detected peptides of suitable length, the peptides listed in said working list (i) are located at positions corresponding to positions in said target peptide that are available for interacting with said antigen recognition molecule, and (ii) each contain at least two amino acids that are identical to corresponding amino acids in said target peptide; e) determining the binding affinity of each of the peptides listed in said working list for said MHC molecule; f) filtering the working list to include only peptides that have a calculated binding affinity for the MHC molecule higher than a first threshold, thereby generating a working list of off-target peptides; g) providing as output a working list of said off-target peptides and / or the number of said off-target peptides in said working list. A non-transitory computer-readable medium containing instructions.

2. The instructions, when executed by the processor, cause the computing device to: estimating the number of peptides in said working list of off-target peptides that are expressed in essential normal tissues; 10. The non-transitory computer-readable medium of claim 1, wherein the non-transitory computer-readable medium provides as output the number of off-target peptides expressed in essential normal tissues.

3. The instructions, when executed by the processor, cause the computing device to: for each peptide in said working list of peptides, determining whether such peptide is expressed in an essential normal tissue; 10. The non-transitory computer readable medium of claim 1, wherein the working list is filtered to include only peptides expressed in essential normal tissues.

4. The instructions, when executed by the processor, cause the computing device to:

4. The non-transitory computer-readable medium of claim 3, wherein the medium generates a list of potential secondary target peptides comprising peptides that have a calculated binding affinity for the MHC molecule that is higher than the first threshold and that have low expression in essential normal tissues.

5. The instructions, when executed by the processor, cause the computing device to: calculating degree of similarity (DoS) scores for the peptides in the working list of peptides, the DoS scores being based at least in part on the number of amino acids identical to amino acids at corresponding positions in the target peptide, the amino acids of the target peptide being available to interact with the antigen recognition molecule; 10. The non-transitory computer-readable medium of claim 1, wherein the working list is filtered to include only peptides having a DoS score above a second threshold.

6. 6. The non-transitory computer-readable medium of claim 5, wherein only positions of the target peptide identified as not bound to the MHC molecule are considered in calculating the DoS score.

7. The instructions, when executed by the processor, cause the computing device to: providing as input a computational representation of the antigen recognition molecule capable of binding to the MHC-target peptide complex; determining the binding affinity of said antigen recognition molecule to a plurality of MHC-peptide complexes each comprising each likely off-target peptide from said working list and said MHC molecule; 10. The non-transitory computer-readable medium of claim 1, wherein the working list is filtered to include only off-target peptides that are likely to contain a binding motif for the antigen recognition molecule.

8. The instructions, when executed by the processor, cause the computing device to: providing as input the expression of off-target peptides in essential normal tissues of a particular patient; 10. The non-transitory computer-readable medium of claim 1, wherein the non-transitory computer-readable medium provides as an output an indication of off-target effects in the patient.

9. A non-transitory computer-readable medium configured to communicate with one or more processors of a computing device, the non-transitory computer-readable medium, when executed by the processors, causing the computing device to: a) receiving as input a computational representation of a target peptide presented in complex with a major histocompatibility complex (MHC) molecule (an MHC-target peptide complex); b) identifying, within the total pool of predicted or detected peptides of suitable length, similar peptides that (i) are located at positions corresponding to positions in the target peptide that are available to interact with an antigen-recognizing molecule, and (ii) contain at least two amino acids that are identical to corresponding amino acids in the target peptide; c) determining the binding affinity of each of said identified similar peptides for said MHC molecule; d) identifying off-target peptides based at least in part on identifying similar peptides having a calculated binding affinity for said MHC molecule that is stronger than a first threshold; e) providing the off-target peptide as an output. A non-transitory computer-readable medium containing instructions.

10. A non-transitory computer-readable medium configured to communicate with one or more processors of a computing device, the non-transitory computer-readable medium, when executed by the processors, causing the computing device to: a) selecting two or more potential target peptides predicted to bind to major histocompatibility complex (MHC) molecules among disease-associated peptides; b) estimating the number of off-target peptides associated with each of said potential target peptides; c) ranking the potential target peptides based at least in part on the number of off-target peptides associated with each of the potential target peptides. A non-transitory computer-readable medium containing instructions.

11. The instructions, when executed by the processor, cause the computing device to: calculating a degree of similarity (DoS) score for each of the off-target peptides, such that the DoS score indicates the similarity between each off-target peptide and the target peptide; 11. The non-transitory computer-readable medium of claim 10, wherein the potential target peptides are ranked based at least in part on the DoS scores of off-target peptides associated with each of the potential target peptides.

12. The instructions, when executed by the processor, cause the computing device to:

12. The non-transitory computer-readable medium of claim 11, wherein the DoS score is calculated based at least in part on the number of amino acids in the off-target peptide that are identical to amino acids at corresponding positions in the target peptide, and the amino acids in the target peptide are available to interact with an antigen-recognizing molecule.

13. 12. The non-transitory computer-readable medium of claim 11, wherein only positions of the target peptide identified as not involved in interaction with the MHC molecule are considered in calculating the DoS score.

14. The instructions, when executed by the processor, cause the computing device to:

12. The non-transitory computer-readable medium of claim 11, wherein the non-transitory computer-readable medium calculates a probability of in vivo toxicity for each potential target peptide based at least in part on the DoS score of the off-target peptide.

15. 15. The non-transitory computer-readable medium of claim 14, wherein the probability of in vivo toxicity for each potential target peptide is based at least in part on the number of highly toxic off-target peptides having a DoS score above a predetermined threshold.

16. 11. The non-transitory computer-readable medium of claim 10, wherein the disease-associated peptides in step (a) are identified based at least in part on a comparison of the expression levels of the corresponding mRNA or protein in diseased tissue and essential normal tissue.

17. The method of claim 16, wherein step b) predicts all amino acid positions within the target peptide that are available to interact with the antigen recognition molecule, i) determining the binding affinity of said target peptide to said MHC molecule; ii) generating a plurality of mutated peptide sequences, each involving a mutation at a respective amino acid position of said target peptide; iii) determining the binding affinity of each mutated peptide of said plurality of mutated peptides to said MHC molecule; iv) predicting the amino acid positions available for interacting with an antigen recognition molecule that recognizes the MHC-target peptide complex based at least in part on a comparison of the binding affinity of each mutated peptide with the binding affinity of the target peptide.

10. The non-transitory computer-readable medium of claim 1, comprising:

18. The instructions, when executed by the processor, further cause the computing device to: a) determining the binding affinity of said target peptide for said MHC molecule; b) generating a plurality of mutated peptide sequences, each involving a mutation at a respective amino acid position of said target peptide; c) determining the binding affinity of each mutated peptide of said plurality of mutated peptides for said MHC molecule; d) predicting the amino acid positions available for interacting with an antigen-recognizing molecule based at least in part on comparing the binding affinity of each mutated peptide with the binding affinity of the target peptide.

19. The instructions, when executed by the processor, further cause the computing device to:

10. The non-transitory computer readable medium of claim 9, wherein off-target peptides are identified based at least in part on expression in essential normal tissues.

20. The instructions, when executed by the processor, further cause the computing device to:

10. The non-transitory computer-readable medium of claim 9, wherein off-target peptides are identified based at least in part on the number of amino acids identical to amino acids at corresponding positions of the target peptide that are available to interact with the antigen recognition molecule.