System and methods of using the same for antigen identification
The nucleic acid-based single-chain polypeptide trimers address the limitations of current HLA-restricted peptide identification methods by providing a scalable and sensitive approach to identify immunotherapy targets, particularly for HLA-C alleles, enhancing clinical trial eligibility and population coverage.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2026-03-26
AI Technical Summary
Current methods for identifying HLA-restricted peptides face challenges such as high material input requirements, difficulty in deconvoluting multi-allele HLAs, limited sensitivity for low-abundance peptides, and decreased reliability for HLA alleles with insufficient training datasets, particularly for HLA-C alleles, leading to biased immunopeptidome data and inequities in clinical trial eligibility.
A composition comprising nucleic acid sequences encoding single-chain polypeptide trimers, each comprising an antigen, β2-microglobulin, and an MHC class I allele, with an IC50 of no greater than about 500 nM, allowing for the identification of immunotherapy targets through the expression of these trimers on HLA/TAP knockout cells.
The solution enables accurate and scalable identification of immunotherapy targets across diverse HLA alleles, improving population coverage and clinical trial eligibility by enhancing the sensitivity and reliability of peptide identification, especially for HLA-C alleles.
Smart Images

Figure US2025047033_26032026_PF_FP_ABST
Abstract
Description
DOCKET NO. STFD-OH-PCT PROVISIONAL PATENTSYSTEM AND METHODS OF USING THE SAME FOR ANTIGEN IDENTIFICATIONCROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 697,446, which was filed September 20, 2024, entitled “System and Methods of Using the Same for Antigen Identification,” and U.S. Provisional Application No. 63 / 744,159, which was filed January 10, 2025, entitled “System and Methods of Using the Same for Antigen Identification;” and U.S. Provisional Application No. 63 / 799,119, which was filed May 2, 2025, entitled “System and Methods of Using the Same for Antigen Identification;” all of which are incorporated herein by reference in their entireties.SEQUENCE LISTING
[0002] The electronic sequence listing filed herewith, titled SlTD-011-PCT_SL.xml, created on September 10, 2025, and having a file size of 1 ,326, 104 bytes is incorporated herein by reference in its entirety.STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH
[0003] This disclosure was made with government support under grant number U54CA260517 awarded by the National Institutes of Health. The government has certain rights in this disclosureFIELD
[0004] The disclosure relates to compositions comprising nucleic acid sequences that, when encoded, are useful for identifying immunotherapy targets. The disclosure also relates to methods of identifying immunotherapy targets, and methods of making compositions for identifying immunotherapy targets.DOCKET NO. STFD-011-PCT PROVISIONAL PATENTBACKGROUND
[0005] Antigen presentation for T-cell receptor (TCR) recognition is pivotal in cellular immunity combatting infections and tumors. Understanding which human leukocyte antigen class I complexes (HLA-I) selects specific antigen peptides from the host proteome holds translational promise for antigen-directed immunotherapy. Cunent methods for identifying HL A -restricted peptides encounter limitations exacerbated by extensive HLA polymorphism.
[0006] C D8 T cells are fundamental in adaptive immunity and can specifically eradicate pathogen- infected cells and cancer ceils without harming healthy cells. This specific targeting relies on the recognition of antigens presented by major histocompatibility complex (MHC) class I (MHC-I) or HLA-I molecules (e.g. HLA-A, -B, and -C) in humans1. In HLA-I antigen presentation, peptide fragments are sampled from the cellular proteome through proteolysis, transported to the endoplasmic reticulum (ER) via the TAP complex, trimmed into 8-10 amino acids when necessary, and loaded onto HLA-I molecules upon high affinity binding. The stable peptide-HLA (pHLA) complex is then transported to the cell surface for subsequent T cell recognition2. The set of peptides presented in HLA class I is termed the immunopeptidome. Mass spectrometry' (MS) is instrumental to determine such immunopeptidome and define fundamental rules of antigen presentation, with recent advances in HLA monoallelic MS approaches enabling systematic characterization of allele-specific peptide repertoire. However, current immunopeptidome data remain highly skewed6, as MS-based approaches still face major challenges including requirement for large sample input, difficulty in deconvoluting multi88allele HLAs, and limited sensitivity in detecting low-abundance but clinically relevant peptides such as those from pathogens and mutated oncoprotein’8. Computational prediction algorithms like NetMHC offer rapid and scalable in silico antigen identification but show decreased reliability for HLA alleles with insufficient training datasets, especially for HLA-C alleles with lower surface expression compared to HLA-A / B alleles9’1"
[0007] Understanding the principles of HLA-I alleles in selecting and presenting certain antigens among the entire host cell proteome can provide insights to advance antigen- directed immunotherapies such as vaccines against infectious disease or cancer. One major challenge is the9DOCKET NO. STFD-011-PCT PROVISIONAL PATENT identification of peptides presented by each HLA-I allele, the most polymorphic region of the human genome3. While current methods such as mass spectrometry (MS)-based HLA immunopeptidomes discovery allow the identification of HLA-presented antigen peptides, multiple challenges remain, including massive material input requirements, deconvolution of multi-allele HLAs, and lower sensitivity for peptides derived from pathogens or mutated proteins4'’. Computational prediction algorithms like NetMHC show decreased reliability' for HLA alleles with insufficient training datasets, especially for HLA-C alleles with lower surface expression compared to HLA-A / B alleles6'7. Existing HLA immunopeptidome data are highly biased and enriched for certain well studied alleles such as HLA-A2, while many HLA alleles more present in non-European ancestry groups including those associated with human disease have received scant biochemical characterization. Such biases also lead to substantial inequities in clinical trial eligibility8. In particular, targeting HLA-C may enable population coverage with a smaller number of alleles, but the paucity of antigen binding data has limited this strategy. Other methods, such as in vitro peptide binding and T cell stimulation assays, have limited throughput due to the requirement and cost of individual peptide synthesis9'11. Considerable efforts have been invested over time in enhancing throughput, as evidenced by the recent advancements in methodologies such as EpiScan and TR.-FR.ET with parallel reading, but scalability remains limited due to inherent difficulties12’13.SUMMARY
[0008] The disclosure relates to a composition comprising a first nucleic acid molecule. The first nucleic molecule comprises a first nucleic acid sequence encoding a single-chain polypeptide (rimer. In some embodiments, the first nucleic acid sequence comprises a first region, second region and third region, each region comprising a nucleic acid sequence encoding a subunit of the trimer. In some embodiments the first region comprises a nucleic acid sequence encoding an antigen or antigenic determinant thereof, the second region comprises a nucleic acid sequence encoding [32 microglobulin or a variant thereof; and the third region comprises a nucleic acid sequence encoding a first MHC class I allele or variant thereof In some embodiments, the antigenDOCKET NO. STFD-OH-PCT PROVISIONAL PATENT or antigenic determinant thereof associates to the first MHC class I allele or variant thereof with an IC50 of no greater than about 500 nM. In some embodiments, each of the nucleic acid sequences encoding the antigen or antigenic determinant thereof, the [32 microglobulin, and the MHC nucleic acid sequence comprise a respective 5’ end and respective 3’ end.
[0009] In some embodiments, the composition further comprises a second nucleic acid molecule comprising a second nucleic acid sequence encoding a second single-chain polypeptide trimer. The single-chain polypeptide trimer encoded by a second nucleic acid sequence comprises a first region, a second region, and a third region In some embodiments, the first region comprises a nucleic acid sequence encoding an antigen or antigenic determinant thereof, the second region comprises a nucleic acid sequence comprising a |32microglobulin or a variant thereof, and the third region comprises nucleic acid sequence encoding a second MHC class I allele or variant thereof. In some embodiments, the second nucleic acid molecule comprises a second nucleic acid sequence encoding an antigen or antigenic determinant thereof that associates to the second MHC class I allele or variant thereof with an IC50 of no greater than about 500 nM In some embodiments, the In some embodiments, the first MHC class I allele is different than the second MHC class I allele. In some embodiments, the composition further comprises a third nucleic acid molecule comprising a third nucleic acid sequence encoding a third single-chain polypeptide trimer. The single-chain polypeptide trimer encoded by the third nucleic acid sequence comprises a first region, a second region, and a third region. In some embodiments, the third nucleic acid sequence comprises a first region comprising a nucleic acid an antigen or antigenic determinant thereof, the second region comprises a nucleic acid sequence encoding a p2 microglobulin or variant thereof, and the third region comprises a nucleic acid encoding a third MHC class I allele or variant thereof. In some embodiments, the antigen or antigenic determinant thereof associates to the third MHC class I allele or variant thereof with an IC50 of no greater than about 500 nM.
[0010] In some embodiments, first MHC class I allele is different than the second class I allele, and the third class I allele is different than the first and second class I alleles. In some embodiments, each of the antigens or antigenic determinants thereof are different across each of the first and second or across the first, second and third nucleic acid molecules In some embodiments, each ofDOCKET NO. STFD-OH-PCT PROVISIONAL PATENT the p2 microglobulins or variants thereof are different across each of the first and second nucleic acid molecules or across the first, second and third nucleic acid molecules.
[0011] In some embodiments, the composition comprises a fourth nucleic acid molecule comprising a fourth nucleic acid sequence encoding a single-chain polypeptide trimer. In some embodiments, the single-chain polypeptide trimer encoded by the fourth nucleic acid sequence comprises a first region, a second region, and a third region. The first region comprises an antigen or antigenic determinant thereof, the second region comprises [32microglobulin or a variant thereof, and the third region comprises a fourth MHC class I allele or variant thereof. The antigen or antigenic determinant thereof is encoded by an antigen nucleic acid sequence within the fourth nucleic acid sequence, the p2microglobulin or a variant thereof is encoded by a [32 nucleic acid sequence within the fourth nucleic acid sequence, and the fourth MHC class I allele or variant thereof is encoded by an MHC nucleic acid sequence within the fourth nucleic acid sequence. In some embodiments, the antigen or antigenic determinant thereof associates to the fourth MHC class ! allele or variant thereof with an IC50 of no greater than about 500 nM. In some embodiments, first MHC class I allele is different than the second class I allele, and the third class I allele is different than the first and second class I alleles, and the fourth MHC class I allele is different than the first, second, and third class I alleles.
[0012] In some embodiments, any of the above summarized compositions comprise a respective antigen or antigenic determinant thereof that has an E-score from about 3 2 to about 5. In some embodiments, the first, second, third, and / or fourth MHC class I alleles are selected from HLA-A, HLA-B, or HLA-C alleles. In some embodiments, the first, second, third, and / or fourth MHC class I alleles are chosen from: A*01 :01, A*02:01, A*02:05, A*02:12, A*03:01, A*l l:01, A*23:01, A *24:02, A*30:01, A.*31 .01, A*31:08, A*34:01, A*33:O3, A*68:01, B*07:02, B*08:01, B*08:02, B*13:()l, B*15:01, B*15:02, B*15:03, B*18:01, B*27:01 , B*27:05, B *27:02, B*35:()l, B*35:02, B*35:08, B*39:06, B*40:01, B*40:06, B*40: 10, B*41 :01, B*44:0, B*46:01, B*48:03, B*50:01, B*51 :01 , B*52:01, B*54:01, B*56:01, B*57:01 , B*58:01, C*03:04, C*04:01, C *06:02, C*07:01, C*07:02, C*08:02, C*12:03. In some embodiments, the first and second; or the first, second and third; or the first, second, third and fourth MHC class I alleles are different MHC Class I allelesDOCKET NO. STFD-OH-PCT PROVISIONAL PATENT and are chosen from: A*()l :01, A*02:0l , A*02:05, A*02: 12, A*03:01, A*ll :01, A*23:0l , A*24:02,A*30:01, A*31:01, A*31:08, A*34:01, A*33:O3, A*68:01, B*07:02, B*08:01, B*08:02, 8*13:01 , 6*15:01, 8* 15:02, B*15:03, B*18:01 , B*27:01, B*27:05, B*27:02, 8*35:01 , B*35:02, 8*35:08, 8*39:06, B*40:01, B*40:06, B*40: 10, 8*41:01, 6*44:0, 6*46:01, 8*48:03, 8*50:01, B*51:01, 8*52:01, B*54:01, B*56:01, B*57;01, 8*58:01 , C*03:04, C*04:01, C*06;02, C*07:01, C*07:02, C*08:02, C* I2:O3.
[0013] The disclosure also relates to a composition comprising a first nucleic acid molecule and a second nucleic acid molecule. In some embodiments, the first nucleic acid molecule comprises a first nucleic acid sequence encoding a first single-chain polypeptide trimer, and the first singlechain polypeptide trimer comprises a first subunit, a second subunit, and a third subunit. In some embodiments, the first subunit comprises an antigen or antigenic determinant thereof, the second subunit comprises a p2microglobulin or a variant thereof, and the third subunit comprises a first MHC class I allele or variant thereof. In some embodiments, the second subunit is free of an endogenous signal sequence. In some embodiments, the third subunit is free of its endogenous MHC signal sequence. In some embodiments, the first subunit comprises a nucleic acid sequence encoding an antigen or antigenic determinant thereof the second region comprises a nucleic acid sequence encoding P2 microglobulin or a variant thereof free of the endogenous microglobulin signal sequence; and the third region comprises a nucleic acid sequence encoding a first MHC class! allele or variant thereof free of its MHC endogenous signal sequence. In some embodiments, the MHC class I allele or variant thereof has a point mutation at amino acid 84.
[0014] In some embodiments, the antigen or antigenic determinant thereof is encoded by an antigen nucleic acid sequence within the first nucleic acid sequence, the p2microglobulin or a variant thereof is encoded by a [32 nucleic acid sequence within the first nucleic acid sequence, and the first MHC class I allele or variant thereof is encoded by an MHC nucleic acid sequence within the first nucleic acid sequence. The second nucleic acid molecule comprises a second nucleic acid sequence encoding a second single-chain polypeptide trimer. The second single-chain polypeptide trimer comprises a first subunit, a second subunit, and a third subunit; wherein the first subunit comprises an antigen or antigenic determinant thereof, the second subunit comprisesDOCKET NO. STFD-011-PCT PROVISIONAL PATENTP2tnicroglobulin or a variant thereof, and the third subunit comprises a second MHC ciass I allele or variant thereof. In some embodiments, the antigen or antigenic determinant thereof is encoded by an antigen nucleic acid sequence within the second nucleic acid sequence, the p2microglobulin or a variant thereof is encoded by a P2 nucleic acid sequence within the second nucleic acid sequence, and the second MHC class I allele or variant thereof is encoded by an MHC nucleic acid sequence within the second nucleic acid sequence. In some embodiments, the first single-chain polypeptide trimer is different than the second single-chain polypeptide trimer, In some embodiments, the composition further comprises at least one of a third nucleic acid molecule or a third and a fourth nucleic acid molecule. In some embodiments, the third nucleic acid molecule comprises a third nucleic acid sequence encoding a third single-chain polypeptide trimer comprising a first subunit, a second subunit, and a third subunit; wherein the first subunit comprises an antigen or antigenic determinant thereof, the second subunit comprises p2microglobulin or a variant thereof, and the third subunit comprises a third MHC class I allele or variant thereof. In some embodiments, the antigen or antigenic determinant thereof is encoded by an antigen nucleic acid sequence within the third nucleic acid sequence, the P2microglobulin or a variant thereof is encoded by a p2 nucleic acid sequence within the third nucleic acid sequence, and the third MHC class 1 allele or variant thereof is encoded by an MHC nucleic acid sequence within the third nucleic acid sequence. In some embodiments, the fourth nucleic acid molecule comprises a fourth nucleic acid sequence encoding a fourth single-chain polypeptide trimer. The fourth single-chain polypeptide comprises a first subunit, a second subunit, and a third subunit; wherein the first subunit comprises an antigen or antigenic determinant thereof the second subunit comprising p2microglobulin or a variant thereof, and the third subunit comprises a fourth MHC class I allele or variant thereof. In some embodiments, the antigen or antigenic determinant thereof is encoded by an antigen nucleic acid sequence within the fourth nucleic acid sequence, the P2microglobulin or a variant thereof is encoded by a p2 nucleic acid sequence within the fourth nucleic acid sequence, and the fourth MHC class I allele or variant thereof is encoded by an MHC nucleic acid sequence within the fourth nucleic acid sequence. In some embodiments, the first single-chain polypeptide trimer is different than the second single-chain polypeptide trimer, theDOCKET NO. STFD-011-PCT PROVISIONAL PATENT third single-chain polypeptide trimer is different than the first and the second single-chain polypeptide trimers, and the fourth single-chain polypeptide trimer is different than the first, second, and third single-chain polypeptide trimers. In some embodiments, at least, one of the respective antigen or antigenic determinant thereof in the first, second, third, or fourth single-chain polypeptide has an ESCORE from about 3.2 to about 5. In some embodiments, the first, second, third, and / or fourth MHC class I alleles are selected from HLA-A, HLA-B, or HLA-C alleles. In some embodiments, the first, second, third, and / or fourth MHC class I alleles are selected from .4 *01: 01 , A* 02 : 01 , A * 02 : 05 , A* 02: 12, A * 03 : 01 , A* 11 :01, A* 23 : 01 , A * 24 : 02, A*30 : 01 , A * 31 : 01 , A*31 :08, A*34:0L A*33:03, A*68:01, B*07:02, B*08:01, B*08:02, B*13:01, B*15:01, B*15:02, B*15:03, B*18:01, B*27:01, B*27:05, B*27:02, B*35:01, B*35:02, B*35:08, B*39:06, B*40:01, B*40:06. B*40: 10, B*41:01, B*44:0, B*46:01, B*48:03, B*50:01, B*51:01, B*52:01, B*54:01, B*56:01 , B*57:01, B*58:01, C*03:04, C*04:01 , C*06:02, C*07:01, C*07:02, C*08:02, C*12:03. In some embodiments, the first MHC class I allele is an HLA-A allele, the second MHC I allele is an HLA-B allele, the third MHC class I allele is chosen from an HLA-B or HLA-C, and the third and fourth MHC class I allele is chosen from a HLA-B or HLA-C.
[0015] The disclosure also relates to a composition comprising one or more of the above described first, second, third, or fourth nucleic acid molecules where one or more of the respective antigen nucleic acid sequences is replaced by an insertion sequence.
[0016] The disclosure also relates to a cell comprising one or more of the above summarized compositions. In some embodiments, the cell is an HLA''TAP knockout cell.
[0017] In some embodiments, the disclosure relates to a kit comprising one or more of the above summarized compositions. In some embodiments, the kit further comprises a cell comprising one or more of: the first nucleic acid molecule, the second nucleic acid molecule, the third nucleic acid molecule, and the fourth nucleic acid molecule. In some embodiments, the cell is an HLA / TAP knockout cell. In some embodiments, the cell is capable of being transformed.
[0018] The disclosure relates to a method of identifying an immunotherapy target. Methods of the disclosure comprise selecting from a population of cells comprising any one or more of the above summarized compositions a subpopulation of cells surface displaying one or more respectiveDOCKET NO. STFD-011-PCT PROVISIONAL PATENT single-chain polypeptide trimers. The method further comprises identifying the respective antigen or antigenic determinant thereof within the respective single-chain polypeptide trimer for at least one cell in the subpopulation. In some embodiments, the step of identifying comprises identifying the respective antigen or antigenic determinant thereof within the respective single-chain polypeptide trimer for each cell in the subpopulation. In some embodiments, the method further comprises calculating an E-score for each identified antigen or antigenic determinant. In some embodiments, the method further comprises selecting an identified antigen or antigenic the immunotherapy target. In some embodiments, the method further comprises selecting an identified antigen or antigenic determinant thereof with a high E-score as the immunotherapy target. In some embodiments, the high E-score is from about 3.2 to about 5. In some embodiments, the method further comprises transfecting cells with one or more of the above summarized compositions to create the1population of cells.
[0019] T he disclosure relates to a method of making a plasmid or a method of identifying an immunotherapy target comprising a step of cloning an antigen nucleic acid sequence into the insertion site of one of the above summarized compositions. In some embodiments, the resultant plasmid comprises an expressible nucleic acid sequence encoding a single-chain trimer that comprises a subunit that is an antigen or antigenic determinant thereof.
[0020] In some embodiments, the method of identifying an immunotherapy target further comprises transfecting cells with a nucleic acid molecule library, wherein a plurality of nucleic acid molecules in the nucleic acid library comprise a nucleic acid sequence encoding a singlechain polypeptide trimer comprising a first subunit, a second subunit, and a third subunit, wherein the first subunit comprises an antigen or antigenic determinant thereof: the second subunit comprises a [32microglobulin or a variant thereof; and the third subunit comprises an MHC class I allele or variant thereof. In some embodiments, the antigen or antigenic determinant thereof is encoded by an antigen nucleic acid sequence within the nucleic acid sequence, and the P2microglobulin or a variant thereof is encoded by a p2 nucleic acid sequence within the nucleic acid sequence. In some embodiments, the first MHC class I allele or variant thereof is encoded by an MHC nucleic acid sequence within the nucleic acid sequence. In some embodiments, each ofQDOCKET NO. STFD-011-PCT PROVISIONAL PATENT the antigen nucleic add sequence, the p2 nucleic acid sequence, and the MHC nucleic acid sequence comprise a respective 5’ end and respective 3’ end. In some embodiments, two or more nucleic acid molecules in the nucleic acid library differ from each other in at least one of the antigen nucleic acid sequences and the MHC nucleic acid sequences. In some embodiments, the MHC nucleic acid sequence in different ones of the two or more nucleic acid molecules encodes an HL.A-A, HLA-B, or HLA-C allele. In some embodiments, the HLA-A, HLA-B, or HLA-C allele are independently and respectively selected from A*01:01, A*02:01, A*02:05, A*02:12, A*03 :01, A* 11 :01 , A*23 :01 , A * 24: 02,B*07:02, B*08:01, B*08:02, B*13:01, B*15:01, B*15:02, B*15:O3, B*18:01, B*27:01, B*27:05, B*27:02, B*35:01, B*35;02, B*35:08, B*39:06, B*40:01, B*40:06, B*40:10, B*41:01, B*44:0, B*46:01. B*48:03, B*50:01, B*51 :01, B*52:01, B*54:01, B*56;01, B* 57:01, B*58:01, C*03:04, C*04:01 , C*06:02, C*07:01, C*07:02, C*08:02, or C*12:03. In some embodiments, a set of the different ones of the tw'O or more nucleic acid molecules encodes each of A*01:01, A*02:01, A*02:05, A*02:12, A*03:01, A* 11 :01, A*23:01, A*24:02, A*30:01, A*31 :01 , A*31:08, A*34:01, A*33:03, A*68:01, B*07:02, B*08:01, B*08:02, 8*13:01, B*15:01, B*15:02, B*15:03, B*18:01, B*27:01, B*27:05, B*27:02, B*35:O1, B*35:02, B*35:08, B*39:06, B*40:01, B*40:06, B*40: 10, 8*41 :01 , 8*44:0, B*46:01, B*48:03, 8*50:01 , 8*51 :01 , 8*52:01 , B*54:01, B*56:01, B*57:01 , B*58:01, C*03:04, C*04:01, C*06:02, C*07:01, C*07:02, C*08:02, and C*12:03.[0021 | Methods of identifying an immunotherapy target may comprise a first and second; first, second and third; or a first, second third and fourth nucleic acid molecule, wherein each of the first, second, third and / or fourth nucleic acid molecules comprises an expressible nucleic acid sequence operably linked to a promoter, wherein one or more of the expressible nucleic acid sequences comprises an antigen nucleotide sequence encoding an antigen or an antigenic determinant thereof, wherein each antigen nucleotide sequence encodes one or more of a viral protein, an oncoprotein, or a non-viral intracellular pathogen protein.
[0022] In some embodiments, methods or compositions comprise cells comprising any one or plurality of disclosed nucleic acid molecules, wherein the cells are HLA / TAP knockout cells.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0023] In some embodiments of the method of identifying an immunotherapy target, one or more of the antigen or antigenic determinant thereof is from about 8 to about 10 amino acids in length.
[0024] The disclosure also relates to a system comprising one or more of the above summarized compositions and HLA / T.AP knockout cells, wherein: (a) the cells comprise the one or more of the above summarized compositions, or (b) the cells are in a first container and the one or more compositions are in at least one second container. The disclosure also relates to a system comprising one or more of the above summarized compositions and one or a plurality of cells, wherein: (a) the one or plurality of cells comprise the one or more of the above summarized compositions; or (b) the one or plurality’ of cells are in a first container and the one or more compositions are in at least a second container.
[0025] The disclosure relates to a method of making one or more of the above summarized compositions, wherein the methods comprise a step of cloning one or more antigen sequences into the insertion site of a composition comprising an insertion site or cloning one or more antigen sequences into a multiple cloning site of the nucleic acid molecules. In some embodiments, the composition comprises a library of antigen sequences within one or a plurality of nucleic acid molecules. In some embodiments, the nucleic acid molecules comprise an expressible nucleic acid sequence encoding one or a plurality of antigen sequences or antigenic determinants thereof.
[0026] The disclosure relates to a composition comprising any antigen or antigenic determinant thereof disclosed herein. The disclosure also relates to a composition comprising any one or combination of single-chain trimer polypeptides disclosed herein. In some embodiments, the antigen or antigenic determinant is an immunotherapy target. In some embodiments, the immunotherapy target was identified by a method herein. In some embodiments, the immunotherapy target has an E-score from about 3.2 to about 5.
[0027] In some embodiments, the disclosure relates to a pharmaceutical composition comprising any antigen or antigenic determinant thereof disclosed herein and a pharmaceutically acceptable carrier. In some embodiments, the antigen or antigenic determinant i s an immunotherapy target. In some embodiments, the immunotherapy target is identified by a method herein. In some embodiments, the immunotherapy target has an E-score from about 3.2 to about 5.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0028] In some embodiments, the disclosure relates to a dual HL A and TAP knock-out cell in a system, composition or kit.
[0029] In some embodiments, the disclosure relates to a composition comprising a dual HLA and TAP knock-out cell. In some embodiments, the cell comprises any one or more nucleic acid disclosed herein. In some embodiments, the cell comprises any one or more amino acid sequence disclosed herein. The disclosure also relates to one or a plurality of an amino acid sequences comprising any one or plurality of single-chain trimers disclosed herein. In some embodiments, the disclosure relates to a composition comprising any one or more amino acid sequence disclosed herein directly or translated from nucleic acid sequence.
[0030] In some embodiments, the disclosure relates to a pharmaceutical composition comprising a therapeutically effective amount of: (i) any one or plurality of amino acid sequences disclosed; and (ii) a pharmaceutically acceptable carrier In some embodiments, the pharmaceutical compositions comprise a therapeutically effective amount of one or a plurality of isolated T cells comprising the amino acid sequences recognizing the epitopes disclosed herein or a salt thereof and a pharmaceutically acceptable carrier. In some embodiments, the T cell is an engineered T cell, or an autologous T cell from a human subject comprising a nucleic acid molecule comprising a nucleic acid sequence encoding one or a plurality of sequences from TABLE C.1 or any sequence that recognizes an epitope disclosed herein. In some embodiments, the T cell is an engineered T cell, or an autologous T cell from a human subject comprising a nucleic acid molecule comprising a nucleic acid sequence encoding one or a plurality of amino acid sequence variants from TABLE C. l or any sequence that recognizes an epitope disclosed herein.BRIEF DESCRIPTION OF TH E DRAWINGS
[0031] The following detailed description of embodiments will be better understood when read in conjunction with the appended drawings. For the purpose of illustration, there are shown in the drawings certain embodiments. It is understood, however, that the embodiments are not limited to the precise arrangements and instrumentalities shown. In the drawings:DOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0032] FIGS 1 A- 1F illustrate that single-chain trimers (SCT) of pMHC class I differentiate presentable peptides from those with no or low affinities. FIG. 1A illustrates a schematic of an assay, where SCT with high affinity peptides displayed on cell surface when transfected, while SCT with no binding to MHC alleles failed to. FIG. IB illustrates flowcytometry' measurement of surface pHLA-A*0201 trimers with high (upper right) and low affinity peptides (lower left) in SCTs. .Antibodies either against B2m or HLA-A2 alleles were used to stain the cells. Blank cells were shown in grey as background signals. FIG 1C illustrates imaging of pHLA SCT fusing with eGFP directly at the cytoplasmic tail of HLA-A2. FIG. ID illustrates representative histograms of cell surface SCTs for human HLA-A*0101, 8*0702, C*0401, and mouse H2-kb alleles. 2 s of high affinity peptides (top two sets of histograms) and 2 of negative peptides (bottom two sets of histograms) were shown per allele. The line on the bottom row, behind or next to the hatched peaks, shows the histogram of signals from blank cells. FIG. IE illustrates a plot of gMFI (geometrical mean of fluorescent intensity') vs IEDB measured binding affinity for SCTs containing Y84C mutation (circles) or of wild-type allele (triangles), n-2-4. FIG. IF illustrates a histogram of B2m staining for cell surface pHLA-A2 presentation. 3 conditions were measured, respectively B2m fused with HLA-A2 only, co-expression of B2m-HLA-A2 with a high affinity peptide pp65 preceded with a signal peptide, and lastly a normal SCT for pp65 peptide. Right side: the bargraph showing the quantification of signals. **** p<0 0001.
[0033] FIGS 2A-2H illustrate that. ESCAPE demonstrates good performance and benchmarks across all HLA class I subtypes. FIG. 2A illustrates a schematic of ESCAPE-seq. A pool of peptides across a wide range of affinities were selected from IEDB, synthesized, cloned to build the SCT library'. Cells were transduced with the vims pool of the peptide library' and were sorted and sequenced. FIG. 2B is a representative histogram of cell surface HLA- A *02 trim er expression on cells transduced with pooled virus. The whole cell population was sorted into 4 evenly divided bins on log-scale as indicated on the histogram. FIG. 2C is a scatter plot of IEDB affinity vs E- score calculated for each peptide. A line at 500nM of affinity and a cutoff line for E-score were drawn as indicated. FIG. 2D illustrates plots of recall rate (predicted positive / all positive) vs bins of different IEDB affinities. ESCAPE performance was compared with NetMHC4 either in B ADOCKET NO. STFD-011-PCT PROVISIONAL PATENT(Binding affinity) mode or EL (Elution) mode FIG. 2E illustrates a plot of ROC (Receiver Operating Characteristics) curve of HLA-A2 allele for evaluating ESCAPE in comparison with NetMHC. The AUCs (area under curve) were noted. FIG 2F is aPRC (Precision-recall curve) plot for HLA-A2 allele in comparison with NetMHC. FIG. 2G illustrates barplots of AUC-ROC for all 4 HLA-I alleles to compare the performance of ESCAPE with NetMHC.FIG. 2H illustrates barplots of AUC-PRC for all 4 HLA-I alleles to compare the performance of ESCAPE with NetMHC.
[0034] FIGS. 3A-3J illustrate that ESCAPE-seq reveals presented peptides in SARS-CoV2 Spike, N protein, and strain variants. FIG 3 A illustrates a schematic of peptide pool generation by tiling across full spike S and N protein, and its strain variant mutations. FIG. 3B is a scatter plot of NetMHC prediction vs E-score for each peptide. Lines of cutoff values were drawn at 2% rank for NetMHC El and 3.2 for E-score. Here the E-score is normalized as described in Methods FIG. 3C illustrates consensus motifs for HLA-A*02 and HLA-B*07 from ESCAPE positive SARS peptides. FIG. 3D illustrates plots of E-score for each peptide with 3 H1..A alleles across the Spike protein (x-axis, schematic draw of Spike at bottom). Darker black circles highlight the known presentable peptides from literature. FIG. 3E illustrates plots of E-score for peptides panning protein N. 3 HLA alleles were drawn in aligned position for each peptide Darker black circles highlight the known presentable peptides from literature. FIG. 3F is a hierarchical heatmap showing the cluster and shared peptides from SARS-spike across HLA alleles of interest here FIG 3G illustrates the same plot as above for protein N. FIG. 3H illustrates a neoantigen plot of SARS spike peptides presented by HLA- A*02 allele, where x-axis is the E-score for each mutational peptide while y-axis value is the E-score of its corresponding wild-type peptide. 3 quadrants containing positive antigen peptides were highlighted FIG. 31 is the same plot as in (3H) for HLA- B*07 allele. FIG. 3J is a bar graph of percentage of peptides in 3 quadrants for all 3 HLA-I alleles The top of each bar show's WT+, the middle shows WT / Mut+, and the bottom show's Mut+.
[0035] FIGS. 4A-4G illustrate that combinatorial pools of both peptides and HLAs enable identification of variability in peptide presentation across multiple HLAs. FIG. 4A illustrates schematics of generating, barcoding, and sequencing for combinatorial ESCAPE approach. HereDOCKET NO. STFD-011-PCT PROVISIONAL PATENT a pool of >1500 oncogene peptides randomly combined with 50 human HLA alleles, effectively generating > 75,000 peptide-HLA pairs in one screening. A pool of 100 known pHLA pairs were spiked in during each experiment. The barcodes representing HLA alleles were built into the 12 AA at 3’ end of B2m gene using synonymous mutations. In the end, the libraries of paired peptides and HLA barcodes were amplified by 3 rounds of PCRfor Illumina-based sequencing platform as outlined in the below List of Peptides Sequences, Primers of Table 5. FIG. 4B is a plot of E-score versus IEDB binding affinity for the spike-in pHLA results. The spike- in pool includes 17 known antigens and 80+ positive and negative peptides selected from IEDB. FIG. 4C is a plot of E-score vs NetHMC prediction for the known antigen peptides in the combinatorial pool. NetMHC4 results in both binding affinity mode (Ba, upper) or elution model (El, lower) were shown. The hatched box in the upper right quadrant highlighted the peptides that NetMHC failed the prediction. FIG. 4D is a heatmap view of the Escape score of 1500 oncogene peptides (y-axis) across 50 HLAs (x axis), highlighting a moderate degree of clustering and lack of commonly presented peptides. FIG. 4.E illustrates the that the rate of presented peptides varies across different HLA alleles. FIG. 4F illustrates a count of peptides that were commonly presented by multiple alleles (x-axis). Inset: zoom in plot of peptides count in the shaded region. FIG. 4G illustrates UMAP of HLA allele clustering based on similarity of each allele’s presentation score across all peptides.
[0036] FIGS. 4H-4M illustrate that combinatorial ESCAPE-seq achieves simultaneous profiling of peptide presentation across diverse HLA alleles in one screen. FIG. 4H, Schematics of combinatorial ESCAPE-seq using existing Mass-spectroscopy (MS) elution peptides. 986 most common peptides from the MS dataset, along with 30 HLA alleles were selected (Nat Biotech 2020). FIG. 41, The percentage of epitopes identified by MS that shows a positive E-score by ESCAPE-seq. FIG. 4J, AUC-ROC metrics comparing ESCAPE-seq and MS results across 30 HLA alleles. Peptides were assigned a value of 1 if detected in 1273 MS and 0 if not. FIG. 4K, Schematics of combinatorial ESCAPE-seq used to screen cancer epitopes. A pool of 1500 oncogene peptides tiling across 92 top cancer mutations and 31 oncogenic fusions, was randomly combined with 50 human HLA alleles. All peptides were 9 amino acids in length, and a pool of 100 known pHLA pairs were spiked into the experiment This generated ~75,000peptide-HLADOCKET NO. STFD-011-PCT PROVISIONAL PATENT pairs in one screening. FIG. 41.,, Plot of E-score versus IEDB binding affinity for the spike-in pHLA results. The spike-in pool includes 17 known antigens and 80+positive and negative peptides selected from IEDB. FIG 4M, Plot of E-score versus NetHMC prediction for the known antigen peptides in the combinatorial pool. NetMHC4 results in both binding affinity mode (Ba, upper) or elution model (El, lower) were shown. The hatched box highlights peptides for which NetMHC failed to make a prediction.
[0037] FIGS 5 A-5F illustrate that ESCAPE nominates high priority cancer neoantigens. FIG 5A is a heatmap of number of presented peptides covering each oncogenic driver mutation. Y-axis shows the cancer point mutation along the HLA-A, B, C alleles in x-axis. As there are 9 candidate peptides tiling over each point mutation, the colormap displays the number of presented antigens out of 9. FIG. 5B is a plot of the percentage of driver mutations that were presented by each HLAs in x-axis. FIG, 5C is a histogram of percentage of driver mutations can be presented by individual human, where we repeatedly sampled 2 alleles from HLA-A, B, C alleles, and calculate the mutation coverage to generate the distribution. FIG. 5D is a plot of number of I ll . A alleles that can present peptides containing oncogenic point mutations on x-axis. Mutations along the x-axis were ordered based on HLA allele count on y axis. FIG. 5E is the same plot as FIG. 5D for fusion mutations FIG. 5F illustrates scatter plots of E-score of peptides with mutation (Mut, x-axis) vs. the E-score of its corresponding wild-type peptides (WT, y- axis) from EGFR T790M mutation (left), KRAS G12D (middle), and FLT3 1)835 Y (right panel). The doted lines marked the threshold of E-score that divided the peptides into 4 groups. The dots pointed to highlight the neoantigen peptide known in the literature. FIG. 5F discloses SEQ ID NOS 460, 459, 785, and 599, respectively, in order of appearance.FIGS. 5G-5N illustrate that combinatorial ESCAPE-seq enables population-wide antigen presentation discovery of cancer neoantigens from driver oncogenes. FIG. 5G, Heat map Shi et al. (YU and CHANG), p. 49of the E-score of 1500 oncogene peptides (y-axis) across 1308 50 HLAs (x-axis), showing moderate clustering and absence of commonly presented peptides. FIG. 5H, Count of peptides that were commonly presented by multiple alleles (x-axis). Inset: zoomed-in view of peptide count in the hatched region. FIG 51, Scatter plot of E-score versus HLA allelesDOCKET NO. STFD-011-PCT PROVISIONAL PATENT count for individual peptides highlighted in gray region in (FIG. 5H) that are presented by multiple HLA alleles (x-axis, n>12). Speckled dots represent peptides included in the MS experiments with monoallelic C*0304 cells, while dark dots indicate peptides detected in MS. FIG. 51 discloses SEQ ID NO: 9. FIG. 5J, Aggregated analysis combining all peptides presented across the same mutation. For each point mutation or fusion breakpoint, nine candidate peptides tiling the mutation were analyzed. The grayscale map indicates the number of presented peptides out of nine. FIG. 5K, Heatmap of number of peptides presented for each oncogenic point mutation (Y-axis) across HLA-A, HLA-B, and HLA-C alleles (x-axis). FIG. 5L, Heatmap similar to (FIG. 5K) but for peptides derived from oncogenic fusions. FIG. 5M, Number of HLA alleles capable of presenting peptides containing oncogenic fusion mutation (y-axis) plotted against mutations (x-axis), ordered by HLA allele count. FIG. 5N, Same as (FIG. 5M) but for oncogenic point mutations.
[0038] FIGS. 6A-6I illustrate the following, FIG. 6A: Histogram of B2m staining of wild-type HEK293T cells and cells with HLA-A, B, C knocked out by CRISPR cas9 RNP (HLA-KO), and HLA-KO HEK293T cells expressing NYESO pHLA-A0201 single-chain trim er (SCT). Bottom: isotype antibody staining of HLA-KO HEK293T cells expressing pHLA-A0201. FIG. 6B: Bar plots for (FIG. IB) quantification, comparing gMFI of b2m staining of HEK293T cells expressing single chain trimer (SCT) of HLA-A*0201 with high affinity (HP) peptides vs. fusions with peptides that have no binding affinity (NP). Two peptides were tested for both HP and NP groups. FIG 6C: T-cell stimulation assay showing that NYESO pHL A- A0201 -expressing cells induce eGFP expression in T cells co-expressing a cognate 1G4 TCR and an NFAT reporter. In contrast, pp65 pHLA-A0201 -expressing cells do not. stimulate T cells. FIG. 6D: Direct staining of OVA- pH-2Kb-expressing cells using anti-SIINFEKL (SEQ ID NO: 1303) pH-2Kb and anti-H-2Kb antibodies. FIG. 6E: Effect of TAP 1 / 2 knock-out. gMFI quantification of HP and NP with A*02, B*07 and C*04 alleles in both HLA KO cells andHLA / TAP dual KO cells. FIG. 6F: Motif analysis of amino acids surrounding the conserved Y84 residue across all HLA-A, -B, and -C alleles. The arrow' indicates the conserved Y84 residue, which is mutated to Y84C in SCT constructs. FIG. 6G: Effect of TAP 1 / 2 knock-out. gMFI quantification of HP and NP with A*02, B*07 and C*04 alleles in both HLA KO cells and HLA / TAP dual KO cells. FIG 6H: Fluorescent image of direct eGFPDOCKET NO. STFD-011-PCT PROVISIONAL PATENT fusion with b2m-A*02 allele (top) and A*02 allele (bottom). FIG 61: Compare the effect of TAP KO in FIG. IF. In either HLA KO cells or HLA / TAP dual KO cells, the gMFI of b2m-A2 fusion with or without co-expression of pp65 peptide were compared. *** p<0.001; **** p<0.0()01.
[0039] FIGS. 7A-7M illustrate the following. FIG. 7A: Representative correlation plot of read count between replicates for ESCAPE screen. HLA-A*02 was used here an example. 4 fractions (BG / background, low, medium and high) were highlighted in different colors. Pearson correlation coefficients were shown. FIG. 7B: Correlation of E-score between replicate of ESCAPE screening using HLA alleles including A* 02, A*01, B*07 and C*04. Pearson correlation coefficients of each were shown in the plots. FIG. 7C Scatter plot of IEDB affinities vs. ESCAPE-seq Escore for wildly pe HLA-A*0201 allele. FIG. 7D Scatter plot of IEDB affinities vs. ESCAPE-seq E-score for HLA-A*0101 allele without Y84C mutation. FIG. 7E: Plot of recall rate of ESCAPE-seq vs. IEDB affinity for HLA- A*0l01 allele. NetMHCJba and NetMFIC el prediction results were shown as a comparison. FIGS. 7F and 7G: Similar plots as FIGS. 7C--7E for HLA-B *0702 allele. FIGS. 7 I— J: Similar plots as FIGS. 7C and 7E for HLA.-C*0401 allele. FIG. 7K: Comparison of Pearson correlation coefficient for ESCAPE- seq, NetMHC_ba and NetMHC_el predictions across 4 HLA alleles examined here. FIG. 7L Similar to FIG. 7K for Spearson correlation coefficient. FIG. 7M HLA-Class 1 Alleles v. Spearman.
[0040] FIGS. 8A-8K illustrate the following. FIG. 8 A: Representative correlation plot of read count between replicates for HLA- A*0201. Pearson correlation coefficients were shown. FIG 8B: Correlation of E-score between biological replicates for SARS antigen screening for HLA- A*0201, HLA-A*0101 and HLA-B *0702 alleles. FIG. 8C: Plot of NetMHC el rank and NetMHC ba vs. E-score for HLA-A*0201. Pearson correlation coefficients were shown in each plot. FIG 8D: The same plot as FIG. 8C for HLA- A*0101 , FIG. 8E: The same plots as FIG 8C for HLA-B*0702. FIG. 8F: Plots of count and percentage of presentable peptides (y-axis) vs E- score cutoffs. 3 HLA alleles are compared. Replicate n=2. FIG. 8G: Bar plot of ROC-AUC of NetMHC with ESCAPE E-scores of peptides from IEDB training database (IEDB) and from SARS-CoV2 proteins (SARS-CoV2) for A*01, A*02 and B*07 3 alleles in x-axis. FIG. 8H: The same plot as FIG 8G for PRC-AUC metrics across the 2 peptide pools. **** notes p-value<DOCKET NO. STFD-011-PCT PROVISIONAL PATENT0.0001. FIG. 81: Scatter plots of E-score of antigen peptides containing variant mutation (Mut, x- axis) vs E-score of its corresponding wild-type peptide (WT, y-axis) for spike- HLA-A*0101 pairs. FIG. 8J: The same plots for protein N peptide-HLA pairs for all 3 alleles. FIG. 8K: Quantification of percentage of presented peptides in each group in S3j .
[0041] FIGS 9A-9P illustrate the following. FIG. 9 A, Representative histogram of B2m staining of HEK-KO cells transduced with the combinatorial virus pool. 4 fractions (termed Background, Low, Medium and High) were sorted based on staining signals on log-scale. FIG. 9B, The correlation of reads (left) and E-score (right) of combinatorial ESCAPE-seq. FIG. 9C, Heatmap showing the normalized read count for spike-in peptide HLA pairs. Here the input pool only- consisted of spike-in peptide paired with HLA group-a. All reads from paired group-b were recombination events from various steps involved in the screen, such as templates witch in virus pool, chimeric reads from PCR amplification and library generation, or during sequencing FIG 9D, The violin plot showed the quantification of read counts in (FIG. 9C). In each violin plot, the top and lower bars represent the mean and median inset: zoom in plot, FIG. 9E, .Normalized distribution of E-score for each HLA-I allele, colored in HLA-A, B, C subtypes. FIG. 9F, Percentage of presented peptides varies across different HLA alleles. FIG. 9G, Intensity traces show the detection of peptides (YIMSDSNYV (SEQ ID NO: 599), LTSTVQLIM (SEQ ID NO: 9) and SSYGQQSSL (SEQ ID NO: 510)) with Mass spectroscopy (MS). FIG. 9H, Plot of the percentage of driver mutations from (FIG. 5K) that were presented by each HLAs in x-axis. FIG 91, Presented peptide percentage for oncogenic fusion peptides. FIG. 91, Plot of presented fusion rate from (FIG. 5L) across HLA-I. FIG. 9K, Histogram of percentage of driver mutations (diagonal hatching rising from left to right) or gene fusion mutations (diagonal hatching descending from left to right) can be presented by diploid human cells, where we repeatedly sampled 2 alleles from HLA-A, B, C alleles 1000 times to calculate the mutation coverage to generate the distribution. FIG. 9L: Frequency of peptides that were commonly presented by different number of HLA-B alleles. Insert: zoom-in of shaded region. FIG. 9M: Similar plot as FIG. 9L for HLA-A and HLA- C alleles. FIG. 90: Scatter plot of number of HLA-B alleles (y-axis) that presented each peptide(dot) vs. the number of HLA-A alleles (x-axis) that also presented this peptide. FIG. 90:DOCKET NO. STFD-011-PCT PROVISIONAL PATENTThe same scatter plot as FIG. 9N for HLA-C vs HLA-A. FIG. 9P: The same scater plot as FIG 9N for HLA-C vs HLA-B.
[0042] FIGS. 10A- 10J: FIG. 10A, Representative correlation plot of read counts between replicates for ESCAPE-seq performed on a pool of peptides from the MS dataset. Pearson correlation coefficients are shown in the plot. FIG. 10B, Scatter plot of calculated E-scores between replicates from ( A). FIG. 10C, E-score distributions for peptides that were not detected (ND) or detected by MS. Examples are shown for the HLA-A*03:01 and HLA-A*02:01 alleles FIG. 10D, Representative ROC curves for 4 alleles (A*02:01, B*07:02, C*03:04, and C*04:01) out of a total of 30 alleles in the pool. FIG. 10E, Across all 30 alleles, comparison of E-scores with NetMHC predictions in both EL (elution) and BA (binding affinity) modes Arrows highlight alleles where E-scores show better AUC-ROC metrics. FIG. 10F, Schematic of ESCAPE-seq screening for peptides of from about 8 to about 12 amino acids in length across 25 HL A alleles FIG. 10G, Scatter plots of read counts between biological replicates, and FIG. 10H, scatter plots of calculated E-scores between biological replicates. Pearson correlation coefficients are shown in each plot. FIG. 101, Bar plot showing the number of presented peptides identified by ESCAPE- seq at different peptide lengths. FIG. 10J, E-score heatmap for peptides generated from the same gene locus with varying lengths per allele. Examples include peptides derived from EGFR T790M presented by HLA-B*52:01 and peptides tiling across KRASG12D presented by HLA-A*68:01. The heatmap illustrates changes in E-scores when an epitope is extended by one amino acid at either the N- or C -terminus. In some cases, peptide presentability remains consistent when an additional amino acid is added to the N-terrninus (left and right panels) or the C-terminus (right panel).
[0043] FIGS 11 A- 1 II illustrate the following. FIG. 11 A: Presented peptide percentage across HLA-A, B, C alleles for mutant peptides containing oncogene point mutation and wild-type peptides (WT) without mutation. FIG. 1 IB: Presented peptide percentage for oncogenic fusion peptides. FIG. 11C: Heatmap as (4d) for gene fusion across HLA alleles (x-axis). FIG. 1 ID: Plot of presented fusion rate from FIG. 11C across HLA-I. FIG. HE: The distribution of percentage of presented gene fusions for diploid human cells. It is generated by repeated sampling 1000 timesDOCKET NO. STFD-011-PCT PROVISIONAL PATENT of 2 alleles of HLA- A, B, and C). FIG. HF: Scater plot of E-score of mutant peptide (x-axts) vs. the E-score of its corresponding wild-type (WT, y-axis) peptide for HLAA*0201 allele. The doted lines marked the threshold of E-score that divided the peptide into 4 groups. FIG. J IF discloses SEQ ID NO: 599. FIG. 11 G: Bar plot of percentage of presented peptides in the WT+, Mut+ and WT+Mut+ (both) groups depicted as in FIG. 1 IF across all HL A alleles examined in this paper. FIG. 11H: Aggregated plot from FIG. I IF with HLA A*0201 as an example. Each dot in the plot corresponds to one mutation, with x-axis shows the count of presented peptides containing that mutation while y-axis the count of presented corresponding wild-type peptides FIG. 111: Bar plot of percentage of presented peptides in the WT+ Mut+ and WT+Mut+ (both) groups as (FIG. 5F) across different oncogene mutations, ranked by percentage of population in Mut+.
[0044] FIGS. 12A-I2E illustrate the following. FIG. I2A, Bar plot showing the percentage of presented peptides in the WT+, Mut+ , and WT+Mut+ (both) groups, as depicted in FIGS 16A and 16B, across all HLA alleles examined in this study. The arrow highlights the bar for HLA- A*23:01 . FIG. 12B, Bar plot showing the percentage of presented peptides in the WT+, Mut+, and WT+Mut+ (both) groups, as depicted in FIG. 16C, across different oncogene mutations, ranked by the percentage of peptides in the Mut+ group. FIG. I2C, Gating strategy for the T-cell stimulation assay. Cells were gated sequentially by FSC / SSC, single cells, live cells, and CI)8+ T cells. Double positive cells for IFN-y and TNF-a were quantified and used in the bar plots shown in FIGS. 16G and 161. FIG. 12D, Scatter plot of E-scores for mutant peptides (x-axis) versus the E-scores of their corresponding wild-type (WT) peptides (y-axis) for the HLA-A*2301allele. Dashed lines indicate E-score thresholds that categorize peptides into four groups. Peptides from the EGFR T790M mutation (pl 7, p22, and p23) are labeled. FIG. 12E,Flow cytometry plot corresponding to Fig 6h, showing stimulation by a peptide pool
[0045] FIGS 13A and 13B illustrate the following. FIG. 13A, each HLA alleles’ number of data in IEDB vs its frequency in population (where world population data was downloaded from IEDB: http: / / tools.iedb.org / population / download / ). FIG. 13B, number of alleles sampled v. Fraction of population.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0046] FIGS 14A and 14B illustrate the following. FIG. 14A, E-score distribution for HLA- A*02:01 and peptide motif patten. FIG. 14B, E-score distribution for HLA-A*02:01based on different E-score bins as illustrated in FIG. 14A.
[0047] FIGS. 15A and 15B illustrate examples of E-score distributions for multiple alleles.
[0048] FIGS 16A-16I illustrate that ESCAPE-seq nominates high priority cancer neoantigens FIG. 16A, Schematic of a scatter quadrant plot displaying E-scores for peptides with mutations (Mut, x-axis) versus their corresponding wild-type peptides (WT, y-axis). Dashed lines indicate the E score thresholds, dividing peptides into four groups. The Mut+ only region, where mutant peptides are presented but their corresponding WT peptides are not, is highlighted with hatched lines. FIG. 16B, Example scatter plot of all point mutations as described in (a) for the HLA-A0201 allele. The reported neoantigen FLT3 D835Y is highlighted along with the Mut+ only region (hatched lines) FIG. 16C, Scatter plots of E-scores for mutant peptides (Mut, x-axis) versus their corresponding WT peptides (WT, y-axis) for three mutations: EGFR T790M (left), KRAS G12D (middle), and FL.T3 D835Y (right). Dots indicated with arrows indicate known neoantigen peptides reported in the literature. FIG. 16C discloses SEQ ID NOS 460, 459, 8, and 599, respectively, in order of appearance. FIG. 16D, Schematic of the T-cell stimulation assay used to assess the immunogenicity of specific peptides (see Methods). T cells isolated from healthy donors were activated and stimulated with peptides, followed by intracellular staining for IFN-y and TNF-ct for flow cytometric analysis. FIG. 16E, E-scores for three EGFR T790M mutant peptides (pl 7, p22, and p23; see Table 8) and their corresponding WT peptides presented by HLA-A2301, plotted and compared. FIG, 16E discloses SEQ ID NOS 9, and 457-458, respectively, in order of appearance FIG. 16F, Flow7cytometric analysis of dual intracellular IFN-y and TNF-a staining in donor!356derived T cells stimulated with either DMSO (control) or peptide pl 7 (an EGFR T790Mmutant peptide). FIG. 16F discloses SEQ ID NO: 9. FIG. 16G, Bar plots showing the percentage of IFN-y and TNF-a double 1358positive T cells stimulated with different peptides or conditions. FIG. 16H, Representative flowcytometry' plots of T cells, as in (f), stimulated with DMSO (control) or fusion peptidepll. FIG. 161, Bar plots of the percentage of IFN-y and TNF-aDOCKET NO. STFD-011-PCT PROVISIONAL PATENT double-positive T cells stimulated with various fusion peptides or conditions. FIG 161 discloses SEQ ID NO: 908.DETAILED DESCRIPTION OF EMBODIMENTS
[0049] C ertain terminology is used in the following description for convenience only and is not limiting. The words “right,” “left,” “top,” and “bottom” designate directions in the drawings to which reference is made. Unless specifically defined otherwise, all technical and scientific terms used herein shall be taken to have the same meaning as commonly understood by one of ordinary' skill in the art (e.g., in cell culture, molecular genetics, RNA and detection thereof, immunology', immunohistochemistry, protein chemistry, and biochemistry') The meaning and scope of the terms should be clear, however, in the event of any latent ambiguity, definitions provided herein take precedent over any dictionary or extrinsic definition. Further, unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular. For example, Singleton et al , Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, NY 1994), provide one skilled in the art with a general guide to many of the terms used in the present application. Additionally, the practice of the present disclosure will employ, unless otherwise indicated, conventional techniques of molecular biology (including recombinant techniques), microbiology, cell biology, and biochemistry', which are within the skill of the art. Such techniques are explained fully in the literature, such as, “Molecular Cloning; A Laboratory Manual,” 2nd edition (Sambrook et al., 1989); “Oligonucleotide Synthesis” (M.J. Gait, ed., 1984), “Animal Cell Culture” (R.I. Freshney, ed., 1987); “Methods in Enzymology” (Academic Press, inc.); “Handbook of Experimental Immunology,”' 4th edition (D.M. Weir & C.C. Blackwell, eds., Blackwell Science Inc , 1987); “Gene Transfer Vectors for Mammalian Cells” (J.M. Miller & M.P. Calos, eds., 1987); “Current Protocols in Molecular Biology” (F.M. Ausubel et al., eds., 1987); and “PCR: The Polymerase Chain Reaction,”' (Mullis et al., eds., 1994).
[0050] The words “a” and “one,” as used in the claims and in the corresponding portions of the specification, are defined as including one or more of the referenced item unless specifically stated otherwise. As used in the present disclosure and claims, the singular forms “a,” “an” and “the”DOCKET NO. STFD-011-PCT PROVISIONAL PATENT include plural forms unless the context dearly dictates otherwise. The phrase “at least one” followed by a list of two or more items, such as “A, B, or C” or “A, B, and C,” means any individual one of A, B or C as well as any combination thereof. Likewise, the phrase “one or more of” followed by a list of two or more items, such as “A, B, or C” or “A, B, and C,” means any individual one of A, B or C as well as any combination thereof. The term “and / or” as used in a phrase such as “A and / or B” herein includes both A and B, A or B, A (alone), and B (alone). Likewise, the term “and / or” as used in a phrase such as “A, B, and / or C” encompasses each of the following embodiments: A, B, and C; A, B, or C; A or C; A or B, B or C; A and C, A and B; B and C; A (alone); B (alone); and C (alone).
[0051] It is understood that wherever embodiments are described herein with the language “comprising” otherwise analogous embodiments described in terms of “consisting of’ and / or “consisting essentially of’ are also provided It is also understood that wherever embodiments are described herein with the language “consisting essentially of” otherwise analogous embodiments described in terms of “consisting of” are also provided
[0052] The term “about” as used herein when referring to a measurable value such as an amount, a temporal duration, and the like, is meant to encompass variations of no more than ±20% from the specified value, as such variations are appropriate to perform the disclosed methods. In some embodiments herein comprise modification of any numerical value herein with ±10%, ±5%„ ±1%, or ±0.1 %. For recitation of numeric ranges herein, each intervening number therebetween with the same degree of precision is explicitly contemplated. For example, for the range of about 6 to about 9, the numbers 7 and 8 are contemplated in addition to 6 and 9, and for the range about 6.0 to about 7.0, the numbers 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6,9, and 7.0 are explicitly contemplated.
[0053] The term “culture vessel” as used herein is defined as any vessel suitable for growing, culturing, cultivating, proliferating, propagating, or otherwise similarly manipulating cells. A culture vessel may also be referred to herein as a “culture insert.” In some embodiments, the culture vessel is made out of biocompatible plastic and / or glass, in some embodiments, the plastic is a thin layer of plastic comprising one or a plurality of pores that allow diffusion of protein, nucleic acid, nutrients (such as heavy metals and hormones) antibiotics, and other cell culture mediumDOCKET NO. STFD-011-PCT PROVISIONAL PATENT components through the pores In some embodiments, the culture vessel is designed to contain a hydrogel or hydrogel matrix and various culture mediums. In some embodiments, the culture vessel consists of or consists essentially of a hydrogel or hydrogel matrix. In some embodiments, the only plastic component of the culture vessel is the components of the culture vessel that make up the side walls and / or bottom of the culture vessel that, separate the volume of a well or zone of cellular growth from a point exterior to the culture vessel. In some embodiments, the disclosure relates the culture vessel comprising one or a plurality of cells comprising a nucleic acid sequence encoding one or more amino acid sequences disclosed herein. In some embodiments, the cells are adherent cells.
[0054] “( Crystal” or “crystalline structure,” as used herein, refer to a solid material, whose constituent atoms, molecules, or ions are arranged in an orderly repeating pattern extending in all three spatial dimensions. The process of forming a crystalline structure from a fluid or from materials dissolved in the fluid is often referred to as “crystallization” or “crystallogenesis .” .Protein crystals are almost always grown in solution. The most common approach is to lower the solubility of its component molecules gradually. Crystal growth in solution is characterized by two steps: nucleation of a microscopic crystallite (possibly having only 100 molecules), followed by growth of that crystallite, ideally to a diffraction-quality crystal.
[0055] “X-ray crystallography,” as used herein, is a method of determining the arrangement of atoms within a crystal, in which a beam of X-rays strikes a crystal and diffracts into many specific directions. From the angles and intensities of these diffracted beams, a crystallographer can produce a three-dimensional picture of the density of electrons within the crystal. From this electron density, the mean positions of the atoms in the crystal can be determined, as well as their chemical bonds, their disorder and various other information, as will be known by those skilled in the art.
[0056] As used herein, the terms “treat,” “treated,” or “treating” can refer to therapeutic treatment and / or prophylactic or preventative measures wherein the object is to prevent or slow down (lessen) an undesired physiological condition, disorder or disease, or obtain beneficial or desired clinical results. For purposes of the embodiments described herein, beneficial or desired clinicalDOCKET NO. STFD-011-PCT PROVISIONAL PATENT results include, but are not limited to, alleviation of symptoms; diminishment of extent of condition, disorder or disease; stabilized (z.e., not worsening) state of condition, disorder or disease, delay in onset or slowing of condition, disorder or disease progression; amelioration of the condition, disorder or disease state or remission (whether partial or total), whether detectable or undetectable; an amelioration of at least one measurable physical parameter, not necessarily discernible by the patient; or enhancement or improvement of condition, disorder or disease. Treatment can also include eliciting a clinically significant response without excessive levels of side effects. Treatment also includes prolonging survival as compared to expected survival if not receiving treatment.
[0057] The term “atomic coordinates,” as used herein, refers to a set of three-dimensional coordinates for atoms within a molecular structure. In some embodiments, atomic coordinates are obtained using X-ray crystallography according to methods well-known to those of ordinarily skill in the art of biophysics. Briefly described, X-ray diffraction patterns can be obtained by diffracting X-rays off a crystal. The diffraction data are used to calculate an electron density map of the unit cell comprising the crystal; said maps are used to establish the positions of the atoms (i.e., the atomic coordinates) within the unit cell. Those skilled in the art understand that a set of structure coordinates determined by X-ray crystallography contains standard errors. In other embodiments, atomic co-ordinates can be obtained using other experimental biophysical structure determination methods that can include electron diffraction (also known as electron crystallography) and nuclear magnetic resonance (NMR) methods. In yet other embodiments, atomic coordinates can be obtained using molecular modeling tools which can be based on one or more of ab initio protein folding algorithms, energy minimization, and homology-based modeling. These techniques are well known to persons of ordinary' skill in the biophysical and bioinformatic arts.
[0058] The term “exposing” as used herein refers to bringing a disclosed first protein and a cell, target receptor, second protein or other biological entity together in direct or indirect contact, in such a manner that the amino acid sequence can engage in a protein-protein interaction. This can occur directly by physical contact between the disclosed amino acid sequence acid and an MHC molecule and / or an antigen or plurality of antigens, receptor or other entity In some embodiments,DOCKET NO. STFD-011-PCT PROVISIONAL PATENT the step of exposing refers to indirect protein-protein interactions, e.g., by interacting with another molecule, co-factor, factor, or protein that induces an activity of another molecule.
[0059] As used herein, the term “kit” refers to a set of components provided for purposes of conducting a method of preparing reagents (e.g., vectors, amino acid sequences, cells, etc.} for identifying an immunotherapy target, conducting a method of identifying an immunotherapy target. In some embodiment, a kit comprises devices or conditions for storage, transport, or delivery of various agents (e.g., oligonucleotides, vectors, antigens or antigenic determinants thereof, enzymes, extracellular matrix components, cells, staining reagents, adjuvants, pharmaceutically acceptable carriers, etc. in appropriate containers) and / or supporting materials (e.g., buffers, media, cells, written instructions for performing the assay etc. ) from one location to another. For example, in some embodiments, kits include one or more enclosures (c.g., boxes) containing relevant reaction reagents and / or supporting materials. As used herein, the term “fragmented kit” refers to a kit comprising two or more separate containers that each contain a subportion of total kit. components Containers may be delivered to an intended recipient together or separately. For example, a first container may contain a petri dish or polysterene plate for use in a cell culture assay, while a second container may contain cells, such as control cells. As another example, the kit may comprise a first container comprising a solid support such as a chip or slide with one or a plurality of ligands with affinities to one or a plurality of amino acid sequences disclosed herein and a second container comprising any one or plurality of reagents necessary for the detection and / or quantification of the association between the amino acid sequences of the disclosure and one or more polypeptides, such as antigen sequences from pathogens or cancer cells. The term “fragmented kit” is intended to encompass kits containing Analyte Specific Reagents (ASR”s) regulated under section 520(e) of the Federal Food, Drug, and Cosmetic Act, but are not limited thereto. Indeed, any deliver) / system comprising two or more separate containers that each contain a sub-portion of total kit components are included in the term “fragmented kit.” In contrast, a “combined kit” refers to a delivery system containing all components in a single container (e.g„ in a single box housing each of the desired components).The term “kit” includes both fragmented and combined kits.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0060] As defined herein, the term “inhibition,” “inhibit,” “inhibiting,” and the like in reference to a protein-inhibitor (e.g., antagonist) interaction means negatively affecting (e.g., decreasing) the activity or function of the protein relative to the activity or function of the protein in the absence of the inhibitor. In some embodiments, inhibition refers to reduction of expression of a certain protein within a certain pathway. In some embodiments, inhibition refers to a reduction in the activity of a signal transduction pathway or signaling pathway. Thus, inhibition includes, at least in part, partially or totally blocking stimulation, decreasing, preventing, or delaying activation, or inactivating, desensitizing, or down-regulating signal transduction through exposure of the amount of an amino acid sequence disclosed herein to a culture, system or cell.
[0061] A s used herein, the term “ligand” or “receptor ligand” means a. molecule that specifically binds to an amino acid, either intracellularly or extracellularly. A ligand may be, without the purpose of being limitative, a protein, a (poly)peptide, a lipid, a small molecule, a protein scaffold, an antibody, an antibody fragment, a nucleic acid, a carbohydrate. A ligand may be synthetic or naturally occurring. The term “ligand” includes a “native ligand” which is a ligand that is an endogenous, natural ligand for a native amino acid. In most cases, a ligand is a “modulator” that increases or decreases an intracellular response when it is in contact with, for example binds to, an amino acid that is expressed in a cell. Examples of ligands that are modulators include agonists, partial agonists, inverse agonists, and antagonists, of which a more detailed description can be found further in the specification. In some embodiments, a ligand of the disclosed amino acid sequence is an MHC molecule and / or a polypeptide that is also an antigen.
[0062] The term “conformation” or “conformational state” of a protein refers generally to the range of structures that a protein may adopt at any instant in time. The skilled artisan recognizes that determinants of conformation or conformational state include a protein’s primary structure as reflected in a protein’s amino acid sequence (including modified amino acids) and the environment surrounding the protein. The conformation or conformational state of a protein also relates to structural features such as protein secondary structures (e.g., a-helix, (J-sheet, among others), tertiary' structure (e.g., the three dimensional folding of a polypeptide chain), and quaternary' structure (e.g., interactions of a polypeptide chain with other protein subunits). Post-translationalDOCKET NO. STFD-011-PCT PROVISIONAL PATENT and other modifications to a polypeptide chain such as ligand binding, phosphorylation, sulfation, glycosylation, or attachments of hydrophobic groups, among others, can influence the conformation of a protein. Furthermore, environmental factors, such as pH, salt concentration, ionic strength, and osmolality of the surrounding solution, and interaction with other proteins and cofactors, among others, can affect protein conformation. The conformational state of a protein may be determined by either functional assay for activity or binding to another molecule or by means of physical methods such as X-ray crystallography, NMR, or spin labeling, among other methods. For a general discussion of protein conformation and conformational states, one is referred to Cantor and Schimmel, Biophysical Chemistry; Part I: The Conformation of Biological. Macromolecules, .W.H. Freeman and Company, 1980, and Creighton, Proteins: Structures and Molecular Properties, W.H. Freeman and Company, 1993. A “specific conformation” or “specific conformational state” is any subset of the range of conformations or conformational states that a protein may' adopt. In some embodiments, the polypeptides of the disclosure are presented in a single-chain trimer, such that, folding of the polypeptide sequence results in three subunits aligned such that MHC Class I peptide associates with the antigen or antigenic determinant thereof.
[0063] The term “hydrogel” as used herein is defined as any water-insoluble, crosslinked, three- dimensional network of polymer chains with the voids between polymer chains filled with or capable of being filled with water. The term “hydrogel matrix” as used herein is defined as any three-dimensional hydrogel construct, system, device, or similar structure. In some embodiments, the hydrogel or hydrogel matrix comprises one or more proteins and / or glycoproteins. In some embodiments, the hydrogel or hydrogel matrix comprises one or more of the following proteins: collagen, gelatin, elastin, titin, laminin, fibronectin, fibrin, keratin, silk fibroin, and any derivatives or combinations thereof. In some embodiments, the hydrogel or hydrogel matrix comprises Matrigel* or vitronectin In some embodiments, the hydrogel or hydrogel matrix can be solidified into various shapes, for example, a bifurcating shape designed to mimic a neuronal tract. In some embodiments, the hydrogel or hydrogel matrix comprises poly (ethylene glycol) dimethacrylate (PEG). In some embodiments, the hydrogel or hydrogel matrix comprises Puramatrix. In some embodiments, the hydrogel or hydrogel matrix comprises glycidyl methacryl ate-dextran (MeDex).DOCKET NO. STFD-011-PCT PROVISIONAL PATENTIn some embodiments, two or more hydrogels or hydrogel matrixes are used simultaneously in a ceil culture vessel. In some embodiments, two or more hydrogels or hydrogel matrixes are used simultaneously in the same cell culture vessel but the hydrogels are separated by a wall that create independently addressable microenvironments in the tissue culture vessel such as wells. In a multiplexed tissue culture vessel it is possible for some embodiments to include any number of aforementioned wells or independently addressable location within the cell culture vessel such that a hydrogel matrix in one well or location is different or the same as the hydrogel matrix in another well or location of the cell culture vessel.
[0064] “ Variant” as used herein with respect to a nucleic acid means a nucleic acid sequence comprising (i) a portion or fragment of a referenced nucleotide sequence; (ii) the complement of a referenced nucleotide sequence or portion thereof; (iii) a nucleic acid sequence that is substantially identical to a referenced nucleic acid or the complement thereof, or (iv) a nucleic acid sequence that hybridizes under stringent conditions to the referenced nucleic acid, complement thereof, or a sequences substantially identical thereto. “Variant” with respect to a peptide or polypeptide that differs in amino acid sequence by the insertion, deletion, truncation, conservative substitution of amino acids, or addition of at least one amino acid as compared to a reference sequence, but the peptide or polypeptide retains at least one biological activity of the reference sequence upon which it is based. Variant may also mean a protein with an amino acid sequence that is substantially identical to a referenced protein with an amino acid sequence that retains at least one biological activity. A conservative substitution of an amino acid, / .<?., replacing an amino acid with a different amino acid of similar properties (e.g., hydrophilicity, degree and distribution of charged regions) is recognized in the art as typically involving a minor change. These minor changes can be identified, in part, by considering the hydropathic index of amino acids, as understood in the art Kyte el al., J. Mol. Biol. 157: 105-132 (1982). The hydropathic index of an amino acid is based on a consideration of its hydrophobicity and charge. It is known in the art that amino acids of similar hydropathic indexes can be substituted and still retain protein function. In an aspect, amino acids having hydropathic indexes of ±2 are substituted. The hydrophilicity of amino acids can also be used to reveal substitutions that would result in proteins retaining biological function. ADOCKET NO. STFD-011-PCT PROVISIONAL PATENT consideration of the hydrophilicity of amino acids in the context of a peptide permits calculation of the greatest local average hydrophilicity of that peptide, a useful measure that has been reported to correlate well with antigenicity and immunogenicity. See U.S. Patent No. 4,554,101, which is incorporated herein by reference as if fully set forth. Substitution of amino acids having similar hydrophilicity values can result in peptides retaining biological activity, for example immunogenicity, as is understood in the art. Substitutions may be performed with amino acids having hydrophilicity values within ±2 of each other. Both the hydrophobicity index and the hydrophilicity value of amino acids are influenced by the particular side chain of that amino acid Consistent with that observation, amino acid substitutions that are compatible with biological function are understood to depend on the relative similarity of the amino acids, and particularly the side chains of those amino acids, as revealed by the hydrophobicity, hydrophilicity, charge, size, and other properties. Nucleic acid molecules or nucleic acid sequences of the disclosure include those that encode amino acid sequences described herein and variants or functional fragments thereof that possess no less than about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 96%, about 97%, about 98%, or about 99% sequence identity with the coding sequences of the foregoing. The term “variant” includes polypeptides conjugated to a natural or non-natural chemical moiety, in some embodiments, the polypeptide comprises a polymer, such as polyethylene glycol, and may be comprised of one or more additional derivitizations of cysteine, lysine, or other residues. In addition, variants of the instant disclosure may comprise a linker or polymer, wherein the amino acid to which the linker or polymer is conjugated may be a non-natural amino acid or may be conjugated to a naturally encoded amino acid utilizing techniques known in the art such as coupling to lysine or cysteine. Polymer modification of polypeptides has been reported. U.S, Patent No. 4,904,584, which is incorporated herein by reference as if fully set forth, discloses PEGylated lysine depleted polypeptides, wherein at least one lysine residue has been deleted or replaced with any other amino acid residue. WO 99 / 67291, which is incorporated herein by reference as if fully set forth, discloses a process for conjugating a protein with PEG, wherein at least one amino acid residue on the protein is deletedDOCKET NO. STFD-011-PCT PROVISIONAL PATENT and the protein is contacted with PEG under conditions sufficient to achieve conjugation to the protein.
[0065] The term variant also includes glycosylated variants, such as but not limited to, variants glycosylated at any amino acid position, N-linked or O-linked glycosylated forms of the polypeptide. In addition, splice variants are also included. The term variant also includes heterodimers, homodimers, heteromultimers, or homomultimers of any one or more polypeptide, protein, carbohydrate, polymer, small molecule, linker, ligand, or other biologically active molecule of any type, linked by chemical means or expressed as a fusion protein, as well as polypeptide variants containing, for example, specific deletions or other modifications yet maintain biological activity.
[0066] In some embodiments, variants further comprise an addition, substitution or deletion that modulates biological activity of the variants. For example, the additions, substitution or deletions may modulate one or more properties or activities of the variant. For example, the additions, substitutions or deletions may modulate affinity for a receptor or binding partner, modulate (including but not limited to, increases or decreases) dimerization, stabilize receptor dimers, modulate the conformation or one or more biological activities of a binding partner such as an antigen or MHC molecule, modulate stability of the polypeptide, modulate cleavage by peptidases or proteases, modulate dose, modulate release or bio-availability, facilitate purification, or improve or alter a particular route of administration. Similarly, variants of the present disclosure may comprise protease cleavage sequences, reactive groups, antibody-binding domains (including but not limited to, FLAG or poly-His) or other affinity based sequences (including but not limited to, FLAG, poly-His, GST, etc.) or linked molecules (including but not limited to, biotin) that improve detection (including but not limited to, GFP), purification or other traits of the polypeptide.
[0067] The “percent identity’" of two polynucleotide or two polypeptide sequences is determined by comparing the sequences using the GAP computer program (a part of the GCG Wisconsin Package, version 10.3 (Accel rys, San Diego, Calif.)) using its default parameters “Identical” or “identity” as used herein in the context of two or more nucleic acids or amino acid sequences, maymean that the sequences have a specified percentage of residues that are the same over a specifiedDOCKET NO. STFD-011-PCT PROVISIONAL PATENT region. The percentage may be calculated by optimally aligning the two sequences, comparing the two sequences over the specified region, determining the number of positions at which the identical residue occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the specified region, and multiplying the result by 100 to yield the percentage of sequence identity. In cases where the two sequences are of different lengths or the alignment produces one or more staggered ends and the specified region of compari son includes only a single sequence, the residues of single sequence are included in the denominator but not. the numerator of the calculation. When comparing DNA and RNA, thymine (T) and uracil (U) may be considered equivalent. Identity' determination may be performed manually or by using a computer sequence algorithm such as BLAST or BLAST 2.0. Briefly, the BLAST algorithm, which stands for Basic Local Alignment Search Tool is suitable for determining sequence similarity. Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information (http: / / www.ncbi. nlm.nih.gov). This algorithm involves first identifying high scoring sequence pair (HSPs) by identifying short words of length Win the query sequence that either match or satisfy some positive-valued threshold score T when aligned with a word of the same length in a database sequence. T is referred to as the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood word hits act as seeds for initiating searches to find HSPs containing them. The word hits are extended in both directions along each sequence for as far as the cumulative alignment, score can be increased. Extension for the word hits in each direction are halted when: 1) the cumulative alignment score falls off by the quantity X from its maximum achieved value: 2) the cumulative score goes to zero or below, due to the accumulation of one or more negative- scoring residue alignments; or 3) the end of either sequence is reached. The Blast algorithm parameters W, T and X determine the sensitivity and speed of the alignment. The Blast program uses as defaults a word length (W) of 11, the BLOSUM62 scoring matrix (see Henikoff et al., Proc. Natl. Acad. Sci. USA, 1992, 89, 10915- 10919, which is incorporated herein by reference in its entirety) alignments (B) of 50, expectation (E) of 10, M=5, N=4, and a comparison of both strands. The BLAST algorithm (Karlin et al., Proc. Natl. Acad. Sci. USA, 1993, 90, 5873-5787, which is incorporated herein by reference in itsDOCKET NO. STFD-011-PCT PROVISIONAL PATENT entirety) and Gapped BLASI' perform a statistical analysis of the similarity between two sequences. One measure of similarity provided by the BLAST algorithm is the smallest sum probability (P(N)), which provides an indication of the probability by which a match between two nucleotide sequences or amino acid sequences would occur by chance. For example, a nucleic acid is considered similar to another if the smallest sum probability in comparison of the test nucleic acid to the other nucleic acid is less than about 1, less than about 0.1, less than about 0.01, and less than about 0.001 Two single-stranded polynucleotides are “the complement” of each other if their sequences can be aligned in an anti-parallel orientation such that every nucleotide in one polynucleotide is opposite its complementary' nucleotide in the other polynucleotide, without the introduction of gaps, and without unpaired nucleotides at the 5’ or the 3’ end of either sequence. A polynucleotide is “complementary” to another polynucleotide if the two polynucleotides can hybridize to one1another under moderately stringent conditions. Thus, a polynucleotide can be complementary to another polynucleotide without being its complement.
[0068] “Nucleic acid” or “oligonucleotide” or “polynucleotide” as used herein may mean at least two nucleotides covalently linked together. The depiction of a single strand also defines the sequence of the complementary' strand. Thus, description of a single strand of nucleic acid also describes the complementary'' strand of a described single strand. In some embodiments, the disclosure of a single strand of nucleic acid encompasses the disclosed strand and the complementary-' strand. Variants of a nucleic acid may be used for the same purpose as a given nucleic acid. Thus, a nucleic acid also encompasses substantially identical nucleic acids and complements thereof. A single strand provides a probe that may hybridize to a target sequence under stringent hybridization conditions. Thus, a nucleic acid also encompasses a probe that hybridizes under stringent hybridization conditions.
[0069] Nucleic acids may be single stranded or double stranded, or may contain portions of both double stranded and single stranded sequence. The nucleic acid may be DNA, both genomic and cDNA, RNA, or a hybrid, where the nucleic acid may contain combinations of deoxyribo- and ribonucleotides, and combinations of bases including uracil, adenine, thymine, cytosine, guanine, inosine, xanthine hypoxanthine, isocytosine and isoguanine. Nucleic acids may be obtained byDOCKET NO. STFD-011-PCT PROVISIONAL PATENT chemical synthesis methods or by recombinant methods. In some embodiments, the nucleic acid is isolated from an organism.
[0070] The term “polypeptide” encompasses two or more naturally or non-naturally-occurring amino acids joined by a covalent bondan amide bond). Polypeptides as described herein include full-length proteins (tyg., fully processed pro-proteins or full-length synthetic polypeptides) as well as shorter amino acid sequences (e.g., fragments of naturally-occurring proteins or synthetic polypeptide fragments)
[0071] The terms “functional fragment” means any portion of a polypeptide or nucleic acid sequence from which the respective full-length polypeptide or nucleic acid relates. In some embodiments, a functional fragment is a portion of a full-length or wiki-type nucleic acid sequence that encodes any one of the nucleic acid sequences disclosed herein, and said portion encodes a polypeptide of a certain length and / or structure that is less than full-length but encodes a domain that still biologically functional as compared to the full-length or wild-type protein. In some embodiments, the functional fragment may have a reduced biological activity, about equivalent biological activity, or an enhanced biological activity as compared to the wild-type or full-length polypeptide sequence upon which the fragment is based.
[0072] The term “salt” refers to acidic salts formed with inorganic and / or organic acids, as well as basic salts formed with inorganic and / or organic bases. Examples of these acids and bases are well known to those skilled in the art. Such acid addition salts will normally be pharmaceutically acceptable although salts of non-pharmaceutically acceptable acids may be of utility in the preparation and purification of the compound in question Acid addition salts of the compounds of the disclosure are most suitably formed from pharmaceutically acceptable acids, and include for example those formed with inorganic acids e.g. hydrochloric, hydrobromic, sulphuric or phosphoric acids and organic acids e.g. succinic, malaeic, acetic or fumaric acid. Other non- pharmaceutically acceptable salts e.g. oxalates can be used for example in the isolation of the compounds of the disclosure, for laboratory use, or for subsequent conversion to a pharmaceutically acceptable acid addition salt. Also included within the scope of the disclosure are solvates and hydrates of the disclosure.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0073] The conversion of a given compound salt to a desired compound salt is achieved by applying standard techniques, in which an aqueous solution of the given salt is treated with a solution of base e.g. sodium carbonate or potassium hydroxide, to liberate the free base which is then extracted into an appropriate solvent, such as ether. The free base is then separated from the aqueous portion, dried, and treated with the requisite acid to give the desired salt
[0074] Examples of salts also include, without limitation, the non-toxic inorganic and organic acid addition salts such as the hydrochloride derived from hydrochloric acid, the hydrobromide derived from hydrobromic acid, the nitrate derived from nitric acid, the perchlorate derived from perchloric- acid, the phosphate derived from phosphoric acid, the sulphate derived from sulphuric acid, the formate derived from formic acid, the acetate derived from acetic acid, the aconate derived from aconitic acid, the ascorbate derived from ascorbic acid, the benzenesulphonate derived from benzensulphonic acid, the benzoate derived from benzoic acid, the cinnamate derived from cinnamic acid, the citrate derived from citric acid, the embonate derived from embonic acid, the enantate derived from enanthic acid, the fumarate derived from fumaric acid, the glutamate derived from glutamic acid, the glycolate derived from glycolic acid, the lactate derived from lactic acid, the maleate derived from maleic acid, the malonate derived from malonic acid, the mandelate derived from mandelic acid, the methanesulphonate derived from methane sulphonic acid, the naphthalene-2-sulphonate derived from naphtalene-2-sulphonic acid, the phthalate derived from phthalic acid, the salicylate derived from salicylic acid, the sorbate derived from sorbic acid, the stearate derived from stearic acid, the succinate derived from succinic acid, the tartrate derived from tartaric acid, the toluene-p-sulphonate derived from p-toluene sulphonic acid, and the like Particularly preferred salts are sodium, lysine and arginine salts of the compounds of the disclosure. Such salts can be formed by procedures well known and described in the art. Other acids such as oxalic acid, which cannot be considered pharmaceutically acceptable, can be useful in the preparation of salts useful as intermediates in obtaining a chemical compound of the disclosure and its pharmaceutically acceptable acid addition salt. Metal salts of a chemical compound of the disclosure include alkali metal salts, such as the sodium salt of a chemical compound of the disclosure containing a carboxy group. Mixtures of isomers obtainable accordingDOCKET NO. STFD-011-PCT PROVISIONAL PATENT to the disclosure can be separated in a manner known per se into the individual isomers, diastereoisomers can be separated, for example, by partitioning between polyphasic solvent mixtures, recrystallization and / or chromatographic separation, for example over silica gel or by, e.g., medium pressure liquid chromatography over a reversed phase column, and racemates can be separated, for example, by the formation of salts with optically pure salt-forming reagents and separation of the mixture of diastereoisomers so obtainable, for example by means of fractional crystallization, or by chromatography over optically active column materials.
[0075] “Substantially complementaiy” as used herein may mean that a first sequence, such as the disclosed amino acid sequences, is at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to the complement of a second sequence over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more nucleotides or amino acids, or that the two sequences hybridize under stringent hybridization conditions.
[0076] “Substantially identical” as used herein may mean that, in respect to a first and a second sequence, a first and second sequence are at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical over a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more nucleotides or amino acids, or with respect to nucleic acids, if the first sequence is substantially complementary' to the complement of the second sequence.
[0077] ,4 “non-naturally encoded amino acid” refers to an amino acid that is not one of the 20 common amino acids or pyrolysine or selenocysteine. Other terms that may be used synonymously with the term “non-naturally encoded amino acid” are “non-natural amino acid,” “unnatural amino acid,” “non-naturally-occurring amino acid,” and variously hyphenated and non-hyphenated versions thereof. The term “non-naturally encoded amino acid” also includes, but is not limited to, amino acids that occur by modification (e.g, post-translational modifications) of a naturally encoded amino acid (including but not limited to, the 20 common amino acids or pyrolysine and selenocysteine) but are not themselves naturally incorporated into a growing polypeptide chain by the translation complex. Examples of such non-naturally-occurring amino acids include, but areDOCKET NO. STFD-011-PCT PROVISIONAL PATENT not limited to, N-acetylglucosaminyl-L -serine N acetylglucosaminyl-L-threonine and O phosphotyrosine. A wide variety of non-naturally encoded amino acids are suitable for use in the present disclosure. Any number of non-naturally encoded amino acids can be introduced into an variant. In general, the introduced non-naturally encoded amino acids are substantially chemically inert toward the 20 common, genetically-encoded amino acids ( / .e., alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, and valine) In some embodiments, the non-naturally encoded amino acids include side chain functional groups that react efficiently and selectively with functional groups not found in the 20 common amino acids (including but not limited to, azido, ketone, aldehyde and aminooxy groups) to form stable conjugates. For example, a variant that includes a non-naturally encoded amino acid containing an azido functional group can be reacted with a polymer (including but not limited to, polyethylene glycol) or, alternatively, a second polypeptide containing an alkyne moiety to form a stable conjugate resulting for the selective reaction of the azide and the alkyne functional groups to form a Huisgen {3+2} cycloaddition product.
[0078] As used herein, a phrase referring to an amino acid sequence encoded by a nucleic acid sequence '‘within” a larger nucleic acid sequence means that the encoding sequence is part of the larger nucleic acid sequence. In some embodiments, the encoding sequence is flanked by nucleotides of the larger nucleic acid sequence on either the 5 ’ or 3 ’ ends of the encoding sequence In some embodiments, the encoding sequence is preceded by nucleotides of the larger nucleic acid sequence prior to the 5’ end of the encoding sequence. In some embodiments, the encoding sequence is followed by nucleotides of the larger nucleic acid sequence after the 3’ end of the encoding sequence. In some embodiments, the encoding sequence starts the larger nucleic acid sequence on the 5’ side thereof. In some embodiments, the encoding sequence ends the larger nucleic acid sequence on the 3’ side thereof.
[0079] The term “MHC class I allele” means an “MHC molecule,” 'which is an amino acid sequence from the family that codes for cell surface protein essential for the adaptive immune system. In some embodiments, the MHC molecule is an amino acid based upon a gene cluster onDOCKET NO. STFD-011-PCT PROVISIONAL PATENT human chromosome 6 responsible for encoding a protein expressed on antigen-presenting cells and responsible for association and presentation of exogenous amino acids to immune effector cells. In some embodiments, the MHC molecule is chosen from a human leukocyte antigen (or “I ILA” ) molecule. In some embodiments, the MHC molecule is an HLA molecule that is encoded by human chromosome 6. In some embodiments, the MHC molecule comprises or is one amino acid sequence set forth in Table A or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about. 95%, at least about 96%, at least about 97%, at (east about 98%, or at least about 99% sequence identity to the amino acid sequence. Some embodiments comprise a human nucleic acid sequence encoding an MHC class I allele or a variant thereof. Some embodiments comprise a nucleic acid sequence encoding an amino acid sequence of an MHC class I allele set forth in Table A or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about. 98%, or at least about 99% sequence identity to the amino acid sequence. Some embodiments comprise a nucleic acid sequence set forth in Table A. 1 or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%. at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto.DOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENT ! i : I|DOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENT ! ' ||DOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0080] For Table A.1, the following Table A.2 provides the definition of the symbols.
[0081] Some embodiments comprise a nucleic sequence of comprising the sequence of any one of SEQ NOS: 172 through 221. Some embodiments comprise a nucleic sequence of comprising the sequence of any one of SEQ NOS: 172 through 221 , or a variant thereof having at least about 70%, at least about 72%, at ieast about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%,DOCKET NO. STFD-011-PCT PROVISIONAL PATENT or at least about 99% sequence identity thereto. Some embodiments comprise a nucleic sequence comprising the sequence of any one of SEQ NOS: 172 through 221, or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto, wherein each symbol M, R, W, S, Y, K, V, H, D, B, or N is selected from one alternative expressed in the respective definition thereof.
[0082] In some embodiments, the system or composition disclosed herein comprises a nucleic acid sequence encoding a single-chain amino acid trimer. In some embodiments, the system or composition comprises an amino acid sequence that comprises the single-chain trimer. In some embodiments, the single-chain trimer comprises a P2microglobulin or a variant thereof. In some embodiments, the P2microglobulin or a variant thereof comprises or is one of the amino acid sequences set forth in Table B or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the amino acid sequence. In some embodiments, the nucleic acid molecule comprises a nucleic acid sequence encoding the p2 microglobulin or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about97%, at least about 98%, or at least about 99% sequence identity to the amino acid sequence.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0083] The disclosure relates to a trimer comprising an antigen or antigenic determinant thereof In some embodiments, the system or composition comprises a nucleic acid molecule, such as a plasmid, wherein the nucleic acid molecule comprises a nucleic acid sequence encoding an antigen or antigenic determinant thereof that comprises or is one of the amino acid sequences set forth in Table C or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the amino acid sequence Some embodiments comprise a nucleic acid sequence encoding the antigen or antigenic determinant thereof. Some embodiments comprise an antigen or antigenic determinant thereof that comprises or is one of the amino acid sequences set forth in Tables C, C.1 , C.2, and C.3 or a variant thereof having at least about 70%, at least about 72%, at least about 75%,DOCKET NO. STFD-011-PCT PROVISIONAL PATENT at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the amino acid sequence. Some embodiments comprise an antigen or antigenic determinant thereof that comprises or is one of the amino acid sequences set forth in Tables 1, 2, or 3 or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the amino acid sequence. See the also below Examples section outlining discovery of antigen or antigenic determinant sequences of Tables C. l, C.2, and C.3. In Tables C. l, C.2, and C.3, the allele(s) in which the antigen or antigenic determinant was be presented are listed, and the count indicates the number of total alleles out of 50 that presented the antigen or antigenic determinant.DOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-OH-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0084] In some embodiments, a single chain trimer herein comprises a signal peptide, a first linker, a p2-microglobulin, wherein the P2 -microglobulin is free of a functional endogenous p2- microglobulin signal peptide, a second linker, and an MHC allele, wherein the MHC allele is free of a functional endogenous MHC allele signal peptide. In some embodiments, the single chain trimer comprises an antigen or antigenic determinant thereof Some embodiments comprise a nucleic acid molecule comprising a nucleic acid sequence encoding a single chain trimer herein. In some embodiments, the single chain trimer is the expression product of expressing the nucleic acid sequence. Some embodiments comprise a plasmid comprising a nucleic acid molecule herein. Some embodiments comprise a composition comprising a nucleic acid sequence herein. Some embodiments comprising a nucleic acid molecule herein. Some embodiments comprises a plasmid herein. Some embodiments comprise a pharmaceutical composition comprising a composition herein and a pharmaceutically acceptable carrier. Some embodiment comprise a method of treating a subject comprising administering a composition or pharmaceutical composition herein to the subject.
[0085] In some embodiments, a nucleic acid molecule herein comprises one or more regulatory sequences. In some embodiments, a plasmid herein comprises one or more regulator} / sequences. In some embodiments the nucleic acid molecule comprises one or more regulatory' sequences operably linked to an expressible nucleic acid sequence, where the expressible nucleic acid sequence comprises a nucleic acid sequence encoding a single chain trimer herein. In some embodiments the plasmid comprises one or more regulatory sequences operably linked to anDOCKET NO. STFD-011-PCT PROVISIONAL PATENT expressible nucleic acid sequence, where the expressible nucleic acid sequence comprises a nucleic acid sequence encoding a single chain trimer herein. In some embodiments, the one or more regulatory sequences comprise a promoter region. In some embodiments, the one or more regulator}' sequences comprise a promoter region operably linked to the expressible nucleic acid sequence, such that upon exposure to an RNA polymerase in a cell, the cell expresses the polypeptide encoded by the expressible nucleic acid sequence. In some embodiments, the one or more regulatory sequences comprise one or more of a promoter, an operator, or an enhancer. In some embodiments the expressible nucleic acid comprises one or more sequences encoding one or more of a CAAT box, a CCAAT box, a pribnow box, a TATA box, a SECIS element, a polyadenylation signal, and A-box, a Z-box, a C-box, and E-box, or a G-box.
[0086] Embodiment of the disclosure also relate to a system, kit or nucleic acid molecule comprising a nucleic acid sequence comprising encoding a MHC molecule disclosed herein and a single chain trimer herein. In some embodiments, the nucleic acid sequence comprises a signal peptide, a first linker, a p2-microglobulin, wherein the p2-microglobu!in is free of a functional endogenous p2-microglobulin signal peptide, a second linker, an MHC allele, and a multiple cloning site operably linked to a regulatory sequence operable on whole expressible portion of the of the nucleic acid sequence; wherein the MHC allele is free of a functional endogenous MHC allele signal peptide; and wherein in an active state, the nucleic acid sequence expresses a single chain trimer peptide that comprises a p2-microglobulin, a MHC domain and an antigenic domain In the context of using the nucleic acid molecule as a screening tool, the nucleic acid molecule comprises a multiple cloning site into which one or a plurality of antigens or antigen determinants can be cloned. In this way, the disclosure relates to a kit, composition or library of disclosed nucleic acid molecules or plurality of disclosed nucleic acid molecules, each nucleic acid molecule comprising an expressible nucleic acid sequence encoding one or a plurality of an antigen or antigenic determinant thereof. In some libraries, the antigen or antigenic determinants thereof are each variable sequences that may carry multiple species of the same antigen or antigen determinants thereof or a plurality of heterogenous nucleic acid sequences encoding multiple antigens or antigenic determinants thereof. In this way, the kits, compositions and nucleic aicd14DOCKET NO. STFD-011-PCT PROVISIONAL PATENT molecules are capable of encoding an amino acid sequence with a first second and third subunit, the first subunit comprising the a p2-microglobulin region, the second subunit comprising the NHC molecule, and the third subunit comprising the antigen; wherein the first, second and third subunits are capable of folding into a
[0087] In some embodiments, a nucleic acid molecule herein comprises an expressible nucleic acid sequence encoding a first subunit and a second subunit. In some embodiments, first subunit comprises one of SEQ ID NOS: 85-93, 163, 169, or 170 or a functional variant thereof comprising at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the respective one of one of SEQ ID NO: 85-93, 163, 169, or 170 and free of an endogenous signal peptide. In some embodiments, the second subunit comprises one of SEQ ID NOS: 11-84, 109-158, or 1302 or a functional variant thereof comprising at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at. least, about 95%, at least, about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the respective one of SEQ ID NOS: 11-84, 109-158, or 1302 and free of an endogenous signal peptide. In some embodiments, the expressible nucleic acid sequence encodes the first subunit, the second subunit, and a third subunit. In some embodiments, the third subunit comprises an antigenic fragment of or one of SEQ ID NOS: 94-102 and 104-108 or a functional variant comprising at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about. 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the respective one of SEQ ID NO: 94-102 and 104- 108 In some embodiments, the expressible nucleic acid sequence encodes a signal peptide, the first subunit, the second subunit, and the third subunit. In some embodiments, the signal peptide comprises SEQ ID NO: 159 or a functional variant comprising at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to SEQ ID NO: 159 In some embodiments, the expressible nucleicDOCKET NO. STFD-011-PCT PROVISIONAL PATENT acid sequence encodes the signal peptide, the first subunit, a first linker, the second subunit, a second linker, and the third subunit. In some embodiments, the first linker comprises SEQ ID NO: 161 or a functional variant comprising at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to SEQ ID NO: 161. In some embodiments, the second linker comprises SEQ ID NO: 165 or 167 or a functional variant comprising at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to SEQ ID NO: 165 or 167 In some embodiments, the expressible nucleic acid sequence encodes the first subunit, the second subunit, and the third subunit and at least one of the signal peptide, the first linker, and the second linker, where the sequence encoding the signal peptide, when present, is before the sequence of the first subunit, the sequence encoding the first linker, when present, is between the sequences encoding the first subunit and the second subunit, and the sequence encoding the second linker, when present, is between the sequences encoding the second subunit and the third subunit. In some embodiments, the expressible nucleic acid sequence encodes the first subunit, the second subunit, and the third subunit and at least two of the signal peptide, the first linker, and the second linker, where the sequence encoding the signal peptide, when present, is before the sequence of the first subunit, the sequence encoding the first linker, when present, is between the sequences encoding the first subunit and the second subunit, and the sequence encoding the second linker, when present, is between the sequences encoding the second subunit and the third subunit. In some embodiments, the expressible nucleic acid sequence encodes the first subunit, the second subunit, and the third subunit and each of the signal peptide, the first linker, and the second linker, where the sequence encoding the signal peptide is before the sequence of the first subunit, the sequence encoding the first linker is between the sequences encoding the first subunit and the second subunit, and the sequence encoding the second linker is between the sequences encoding the second subunit and the third subunit.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0088] In some embodiments, a single chain trimer herein comprises a signal peptide, a first linker, a p2-microglobulin lacking an endogenous p2 -microglobulin signal peptide, a second linker, and an MHC allele lacking an endogenous MHC allele signal peptide. In some embodiments, the single chain trimer comprises an antigen or antigenic determinant thereof. Some embodiments comprise a nucleic acid molecule comprising a nucleic acid sequence encoding a single chain trimer herein. In some embodiments, the single chain trimer is the expression product of expressing the nucleic acid sequence. Some embodiments comprise a plasmid comprising a nucleic acid molecule herein
[0089] Some embodiments comprise a composition comprising a nucleic acid sequence herein. Some embodiments comprise a nucleic acid molecule herein Some embodiments comprise a plasmid herein comprising an expressible nucleic acid sequence encoding one or more disclosed amino acid sequences, identified by sequence identifiers or functional variants thereof comprising at least about 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or about 99% sequence identity ot the amino acid sequence identified by sequence identifier. In some embodiments, such plasmid are free of a beta-2 microglobulin sequence. Some embodiments comprise a pharmaceutical composition comprising a composition herein and a pharmaceutically acceptable carrier. Some embodiments comprise a method of treating a subject comprising administering a composition or pharmaceutical composition herein to the subject As used herein, '‘endogenous p2-microglobulin signal peptide ” refers to a signal peptide naturally within a P2-microglobulin. As used herein, “endogenous MHC allele signal peptide” refers to a signal peptide naturally within an MHC allele amino acid sequence.
[0090] In some embodiments, the p2-microglobulin lacking an endogenous p2-microglobulin signal peptide is free of MSRSVALAVLALLSLSGLEG (SEQ ID NO: 222), MGKAAAVVLVTLVALLGL (SEQ ID NO: 223), MKFVLCLAALAVVSCSDNPK (SEQ ID NO: 224), MARVVALVLLGLLSLTGLEA (SEQ ID NO: 225), MSRSVALAVLALLSLSGLEA (SEQ ID NO: 226), MKIALVLLSLLALTLAESN (SEQ ID NO: 227), MSRSVALAVLALLSLSGLEA (SEQ ID NO: 228), MARSVTLVFLVLVSLTGLYA (SEQ ID NO: 229), or MARSVTVIFLVLVSLAVVLA (SEQ ID NO: 230), which are the N-terminal sequences of SEQ ID NOS: 85, 86, 87, 88, 89, 90, 91, 92, and 92, respectively.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0091] In some embodiments, the MHC allele lacking an endogenous MHC signal peptide is free of a signal peptide sequence of Table D, below.DOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTsequence. In some embodiments, the p2-microglobulin variant amino acid sequence is a biologically active fragment of a p2-microglobulin amino acid sequence. In some embodiments, the biologically active fragment of a p2-microglobulin amino acid sequence comprises about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100, about 101, about 102, about 103, about 104, about 105, about 106, about 107, about 108, about 109, about 110, about 111, about 112, about 113, about 114, about 115, about 116, or about 117 amino acids in length. In some embodiments, the biologically active fragment of a 02-microglobulin amino acidDOCKET NO. STFD-011-PCT PROVISIONAL PATENT sequence comprises about 50 amino acids in length to about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100, about 101, about 102, about 103, about 104, about 105, about 106, about 107, about 108, about 109, about 110, about 111, about 112, about 113, about 114, about 115, about 116, or about 117 amino acids in length In some embodiments, the biologically active fragment of a p2-microglobulin amino acid sequence has a length of about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, or about 99% of the length of a reference p2-microglobulin amino acid sequence. In some embodiments, the biologically active fragment of a p2-microblobulin has at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the portion of a reference P2 -microglobulin amino acid sequence corresponding to the fragment. In some embodiments, the reference p2-microglobulin amino acid sequence is one of the p2-microglobulin amino acid sequences herein. In some embodiments, the reference p2-microglobulin amino acid sequence is one of the p2-microglobulin amino acid sequences in Table B or one of SEQ ID NOS: 163, 169, or 170. Some embodiments comprise a nucleic acid encoding a biologically active fragment of a p2-microglobulin amino acid sequence herein.
[0093] In some embodiments, the MHC allele amino acid sequence is a variant amino acid sequence. In some embodiments, the MHC allele variant amino acid sequence is a biologically active fragment of an MHC allele amino acid sequence. In some embodiments, the biologically active fragment of an MHC allele amino acid sequence comprises about 50, about 60, about 70, about 80, about 90, about 100, about 110, about 120, about 130, about 140, about 150, about 160, about 170, about 180, about 190, about 200, about 210, about 220, about 230, about 240, about 250, about 260, about 270, about 280, about 290, about 300, about 310, about 320, about 330, about 335, about 340, about 345, about 350, or about 360 amino acids in length. In some embodiments, the biologically active fragment of an MHC allele amino acid sequence comprisesDOCKET NO. STFD-OH-PCT PROVISIONAL PATENT from about 50 amino acids in length to about 60, about 70, about 80, about 90, about 100, about 110, about 120, about 130, about 140, about 150, about 160, about 170, about 180, about 190, about 200, about 210, about 220, about 230, about 240, about 250, about 260, about 270, about 280, about 290, about 300, about 310, about 320, about 330, about 335, about 340, about 345, about 350, or about 360 amino acids in length. In some embodiments, the biologically active fragment of an MHC allele amino acid sequence has a length of about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, or about 99% of the length of a reference MHC allele amino acid sequence. In some embodiments, the biologically active fragment of an MHC allele has at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the portion of a reference MHC allele amino acid sequence corresponding to the fragment. In some embodiments, the reference MHC allele sequence is one of the MHC allele amino acid sequences herein. In some embodiments, the reference MHC allele amino acid sequence is one of the MHC allele amino acid sequences in Table A. Some embodiments comprise a nucleic acid encoding a biologically active fragment of an MHC allele amino acid sequence herein, in some embodiments, the nucleic acid encoding a biologically active fragment of an MHC allele amino acid sequence comprises about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, or about 99% of the length of a reference nucleic acid encoding an MHC allele amino acid sequence. In some embodiments, the reference nucleic acid encoding an MHC allele amino acid sequence is one in Table A.1.
[0094] In some embodiments, at least one of the signal peptide, the first linker, the second linker, or the antigen or antigenic determinate thereof is a biologically active fragment of a corresponding full length signal peptide first linker, second linker, or antigen or antigenic determinant thereof herein. In some embodiments, the biologically active fragment of a corresponding full length signal peptide first linker, second linker, or antigen or antigenic determinant thereof hereinDOCKET NO. STFD-011-PCT PROVISIONAL PATENT comprises a length of about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about.96%, about 97%, about 98%, or about 99% of the respective full length amino acid sequence In some embodiments, the biologically active fragment of a corresponding full length signal peptide first linker, second linker, or antigen or antigenic determinant thereof herein comprises a length of from about 10% to about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, or about 99% of the respective full length amino acid sequence.
[0095] In some embodiments, the signal peptide comprises a common signal peptide sequence In some embodiments, the signal peptide comprises a human growth hormone signal peptide sequence, a kappa chain signal peptide sequence, an IL-2 signal peptide sequence, or an HLA allele signal peptide sequence. In some embodiments, the signal peptide comprises a human growth hormone signal peptide sequence comprises MATGSRTSLLLAFGLLCLPWLQEGSA (SEQ EDNO: 159) or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. Some embodiments comprise a nucleic acid molecule comprising a nucleic acid sequence encoding MATGSRTSLLLAFGLLCLPWLQEGSA (SEQ ID NO: 159) or a variant thereof having at least about 70%, at least about 72%, at. least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. Some embodiments comprise a nucleic acid molecule comprising a nucleic acid sequence encoding the signal peptide and comprising atggcgacgggttcaagaacttccctactcttgcattggcctgctttgttgccgtggtacaggagggctcggca (SEQ ID NO: 160) or a variant thereof having at least about 70%, at. least about 72%, at. least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0096] In some embodiments, the first linker comprises an amino acid sequence of GCGGSGGGGSGGGGSGG (SEQ ID NO: 161) or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. Some embodiments comprise a nucleic acid molecule comprising a nucleic acid sequence encoding GCGGSGGGGSGGGGSGG (SEQ ID NO: 161) or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. Some embodiments comprise a nucleic acid molecule comprising a nucleic acid sequence encoding the first linker and comprising ggatgcggagggtccggaggtggtggtagcggtggtggaggaagcggagga (SEQ ID NO: 162) or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto.
[0097] In some embodiments, the [32-microglobulin lacking a [32-microglobulin signal peptide comprises an amino acid sequence comprisingIQRTPKIQVYSRHPAENGKSNFLNCYVSGFHPSDIEWLLKNGERIEKXTHSDLSFSKDWS FYLLYYTEFTPTEKDEYACRVNHVTLSQPKIVKWDRDM (SEQ ID NO: 163) or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. Some embodiments comprise a nucleic acid molecule comprising a nucleic acid sequence encoding IQRTPKIQVYSRHPAENGKSNFLNCYVSGFHPSDIEVDLLKNGERIEKVEHSDLSFSKDWS FYLLYYTEFTPTEKDEYACRVNHVTLSQPKIVKWDRDM (SEQ ID NO: 163) or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto Some15DOCKET NO. STFD-011-PCT PROVISIONAL PATENT embodiments comprise a nucleic acid molecule comprising a nucleic acid sequence encoding a P2-microglobulin lacking a p2-microglobulin signal peptide and comprising atccagcgtactccaaagatcaggtttactcacgtcatccagcagagaatggaaagtcaaatttcctgaattgctatgtgtctgggtttcatcc atccgacattgaagttgacttactgaagaatggagagagaattgaaaaagtggagcattcagacttgtctttcagcaaggactggtctttctatc icttgtactacactgaattcacccccactgaaaaagatgagtatgcctgc cgtGTCAACCACGTCACTCTGAGTCAAC CCaagatcgtcaagtgggatcgagacatg (SEQ ID NO: 164) or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. SEQ ID NO: 163 is a human p2-microglobulin.
[0098] In some embodiments, the p2-microglobulin lacking a 32-microglobulin signal peptide comprises an amino acid sequence comprisingSIQKTPQIQVYSRHPPENGKPNILNCYVTQFHPPHIEIQMEKNGKKIPKVEMSDMSFSKD WSFYILAHTEFTPTETDTYACRVKHASMAEPKT'VY'WDRDM (SEQ ID NO: 169) or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. Some embodiments comprise a nucleic acid molecule comprising a nucleic acid sequence encoding SIQKTPQIQVYSRHPPENGKPNILNCYVTQFHPPHIEIQMLKNGKKIPKVEMSDMSFSKD WSFA4I>AHTEFTPTETDTYACRVKHASMAEPK'rVYWDRDK4 (SEQ ID NO: 169) or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. SEQ ID NO: 169 is a Mus musculus p2-microglobulin.
[0099] In some embodiments, the p2-microglobulin lacking a p2-microglobulin signal peptide comprises an amino acid sequence comprisingIQRTPKIQV Y S RHPP ENGKPNFLN C Y VSGFHP S DIE VDL LKNGEKMG K VEHSDLSF SK DW SFYLLYYTEFTPNEKDEYACRVNHVTLSGPRTVKWDRDM (SEQ ID NO: 170) or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, atDOCKET NO. STFD-011-PCT PROVISIONAL PATENT least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. Some embodiments comprise a nucleic acid molecule comprising a nucleic acid sequence encoding IQRTPKIQVYSRHPPENGKPNFLNC'i'VSGFHPSDIEVDLLKNGEKMGKVEHSDLSFSKDW SFYLLYYTEFTPNEKDEYACRVNHVTLSGPRTVKWDRDM (SEQ ID NO: 170) or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. SEQ ID NO: 170 is a niaca p>2-microglobulin.
[0100] In some embodiments, the p2-microglobulin, preferably lacking a p2-microglobulin signal peptide, comprises an amino acid sequence of Table B, or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about98%, or at least about 99% sequence identity thereto. Some embodiments comprise a nucleic acid molecule comprising a nucleic acid sequence encoding an amino acid sequence of Table B, or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto.
[0101] In some embodiments, the nucleic acid encoding a p2-microglobulin comprises a barcodeIn some embodiments, the barcode comprises GTCAACCACGTCACTCTGAGTCAACCC (SEQID NO: 171 ). In some embodiments, the barcode i s oneGTAAACCATGTGACCCTGTCTCAGCCGAAAATCGTT (SEQ ID NO:GTAAACCACGTAACACTTTCTCAACCTAAGATAGTA (SEQ ID NO: 248),GTTAACCATGTGACGCTCAGCCAACCAAAGATTGTC (SEQ ID NO: 249),GTCAACCACGTAACATTGAGCCAACCAAAGATCGTT (SEQ ID NO: 250),GTCAACCATGTAACTCTGAGTCAACCAAAGATCGTG (SEQ ID NO:GTCAACCACGTCACTCTGAGTCAACCCAAGATCGTC (SEQ ID NO:GTCAACCACGTTACACTGAGTCAACCAAAGATTGTA (SEQ ID NO:DOCKET NO. STFD-011-PCT PROVISIONAL PATENTGTAAACCACGTAACCTTGTCCCAACCGAAGATTGTT (SEQ ID NO: 254),GTTAACCATGTTACCCTTAGCCAACCGAAAATCGTA (SEQ ID NO: 255),GTGAATCACGTTACACTGAGTCAGCCAAAGATAGTT (SEQ ID NO: 256),GTTAATCACGTTACGTTGTCACAACCAAAGATCGTG (SEQ ID NO: 257),GTCAACCATGTCACTCTCTCTCAACCCAAGATTGTG (SEQ ID NO: 258),GTCAACCACGTAACGTTGAGTCAGCCAAAGATAGTT (SEQ ID NO: 259),GTGAATCACGTAACTCTTTCACAACCAAAGATTGTT (SEQ ID NO: 260),GTAA ACC’ ACGTTAC ACTCAGCC AACCC A AGAIT'GTG (SEQ ID NO : 261),GTGAACCACGTTACCCTCAGCCAGCCTAAGATCGTG (SEQ ID NO: 262),GTlAArCACGTGACACITl’CACAACClAACMrTGTT (SEQ ID NO: 263),GTAAATCATGTCACGTTGTCCCAACCTAAGATCGTG (SEQ ID NO: 264),GTTAATCATGTTACTTTGAGTCAACCTAAGATTGTC (SEQ ID NO: 265),GTAAACCACGTTACTCTCTCTCAGCCCAAAATCGTA (SEQ ID NO: 266),GTTAACCATGTGACTTTGTCCCAGCCAAAGATTGTA (SEQ ID NO: 267),GTGAATCACGTTACGTTGTCTCAACCAAAGATAGTA (SEQ ID NO: 268),GTTAACCATGTCACTCTGTCCCAGCCGAAGATTGTT (SEQ ID NO: 269),GTGAACCACGTGACGCTTTCACAACCTAAGATCGTA (SEQ ID NO: 270),GTAAATCACGTGACGCTTAGTCAACCTAAGATAGTA (SEQ ID NO: 271),GTTAACCATGTTACACTCAGTCAACCAAAGATTGTG (SEQ ID NO: 272),GTGAACCATGTAACACTTAGTCAGCCAAAGATTGTC (SEQ ID NO: 273),GTGAATCATGTCACGTTGAGTCAGCCCAAAATCGTG (SEQ ID NO: 274),GTGAACCATGTCACTCTTTCACAGCCTAAGATAGTT (SEQ ID NO: 275),GTCAACCACGTTACGCTGAGCCAACCAAAGATCGTT (SEQ ID NO: 276),GTCAACCACGTTACACTTTCTCAGCCCAAGATAGTA (SEQ ID NO: 277),GTCAATCACGTCACTCTTTCCCAGCCGAAGATAGTG (SEQ ID NO: 278),GTGAATCATGTTACTCTGAGTCAGCCTAAGATAGTC (SEQ ID NO: 279),GTCAACCACGTTACTTTGTCCCAGCCGAAGATTGTG (SEQ ID NO: 280),GTAAACCACGTAACTTTGAGCCAACCCAAGATTGTA (SEQ ID NO: 281),DOCKET NO. STFD-011-PCT PROVISIONAL PATENTGTGAATCATGTAACTCTTTCCCAACCAAAGATTG'rG (SEQ ID NO: 282),GTTAACCACGTAACTCTTAGTCAACCAAAGATAGTT (SEQ ID NO: 283),GTAAATCATGTCACTCTTAGTCAACCAAAGATTGTT (SEQ ID NO: 284),GTGAATCACGTGACGCTTTCTCAGCCAAAGATTGTT (SEQ ID NO: 285),GTTAATCATGTAACCCTCTCCCAACCCAAAATCGTT (SEQ ID NO: 286),GTAAATCACGTGACACTCAGCCAACCCAAGATCGTT (SEQ ID NO: 287),GTCAATCATGTCACGCTCTCCCAGCCTAAGATAGTC (SEQ ID NO: 288),GTAAAl’CATGTGACCCTTAGTCAACCCAAGATTGTT (SEQ ID NO:GTCAACCATGTTACACTCAGTCAACCTAAGATAGTA (SEQ ID NO: 290),GTAAATCATGTCACCTTGAGTCAACCCAAGATAGTT (SEQ ID NO: 291),GTTAATCATGTTACCCTTTCTCAACCTAAGATAGTG (SEQ ID NO: 292),GTTAATCATGTCACCCTCAGTCAACCGAAAATCGTA (SEQ ID NO: 293),GTTAACCATGTCACCCTCTCACAGCCTAAGATTGTG (SEQ ID NO: 294),GTCAACCATGTGACGCTCAGCCAGCCCAAGATAGTT (SEQ ID NO: 295), orGTAAACCATGTCACGTTGAGTCAGCCCAAAATCGTT (SEQ ID NO: 296). In some embodiments, the each barcode of SEQ ID NOS: 247-296 is a bar code for a nucleic acid sequence encoding each of SEQ ID NOS: 109-158, respectively. In some embodiments, the each barcode of SEQ ID NOS: 247-296 is a barcode of SEQ ID NOS: 172-221, respectively.|0102| In some embodiments, the second linker comprises an amino acid sequence of GGGSGGGSGGGSHIRNGGGSGGGSGGS (SEQ ID NO: 165) or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. Some embodiments comprise a nucleic acid molecule comprising a nucleic acid sequence encoding GGGSGGGSGGGSHIRNGGGSGGGSGGS (SEQ ID NO: 165) or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. Some embodiments comprise aDOCKET NO. STFD-011-PCT PROVISIONAL PATENT nucleic acid molecule comprising a nucleic acid sequence encoding a second linker and comprising ggtggtggctctggtggaggcagtggaggaggttCCCATATAAGAAAcggaggaggtagtggcggtgggagtggcggatcc (SEQ ID NO: 166) or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto.
[0103] In some embodiments, the second linker comprises an amino acid sequence of GGGGSGGGGSGGGGSGGGGSGS (SEQ ID NO: 167) or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. Some embodiments comprise a nucleic acid molecule comprising a nucleic acid sequence encoding GGGGSGGGGSGGGGSGGGGSGS (SEQ ID NO: 167) or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96 %, at least about 97%, at least about 98 %, or at least about 99% sequence identity thereto. Some embodiments comprise a nucleic acid molecule comprising a nucleic acid sequence encoding a second linker and comprising ggaggtggtggcagcggtggagggggctccggcggtgggggatccggcggcggcggaagcggctct (SEQ ID NO: 168) or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto.
[0104] In some embodiments, the MHC allele, preferably without an MHC allele signal peptide, comprises an amino acid sequence of Table A, or a variant thereof having at least about 70%, at least about 72%, at least, about 75%, at least, about 80%, at. least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. Some embodiments comprises a nucleic acid molecule comprising a nucleic acid sequence encoding an amino acid sequence of Table A, or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least aboutDOCKET NO. STFD-OH-PCT PROVISIONAL PATENT80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto.
[0105] In some embodiments, the antigen or antigenic determinant thereof comprises an amino acid sequence of Table C, or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. Some embodiments comprises a nucleic acid molecule comprising a nucleic acid sequence encoding an amino acid sequence of Table C, or a variant thereof having at least about 70%, at least about 72%, at least about 75%, at least about 80%, at least about 85%, at least about 87%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto.
[0106] Some1embodiments herein comprise1a p2-microblobulin amino acid sequence or nucleic acid encoding the same, or a variant of either, with the native signal peptide (or sequence encoding the same) lacking. In some embodiments, the p2-microblobulin amino acid sequence comprises a barcode sequence. Some embodiments of a barcode sequence are shown in the above P2- microglobulin sequences of SEQ ID NOS: Some embodiments herein comprise an HLA amino acid sequence or nucleic acid encoding the same, or a variant of either, with the native signal peptide (or sequence encoding the same) lacking. Some embodiments herein comprise an HLA amino acid sequence comprising a Y84C mutation. Some embodiments comprise an HLA amino acid sequence herein but comprising a Y84C mutation if not already present. Some embodiments comprise an nucleic acid sequence encoding an HLA amino acid sequence herein but encoding a Y84C mutation if not already present.
[0107] Some embodiments comprise a composition The composition comprises a first nucleic acid molecule. The first nucleic molecule comprises a first nucleic acid sequence encoding a p2~ microglobulin lacking a p2-microglobulin signal peptide single-chain polypeptide trimer. The single-chain polypeptide trimer comprises a first subunit, a second subunit, and a third subunit, wherein the first subunit comprises an antigen or antigenic determinant thereof, wherein the second subunit comprises P2microglobu!in or a variant thereof, and the third subunit comprises a firstDOCKET NO. STFD-OH-PCT PROVISIONAL PATENTMHC class I allele or variant thereof The antigen or antigenic determinant thereof is encoded by an antigen nucleic acid sequence within the first nucleic acid sequence. The p2microglobulin or a variant thereof is encoded by a p2 nucleic acid sequence within the first nucleic acid sequence The first MHC class I allele or variant thereof is encoded by an MHC nucleic acid sequence within the first nucleic acid sequence Each of the antigen nucleic acid sequence, the p2 nucleic acid sequence, and the MHC nucleic acid sequence comprises a respective 5’ end and respective 3’ end. In some embodiments, the antigen or antigenic determinant thereof associates to the first MHC class I allele or variant thereof with an IC50 of no greater than about 500 nM.[OIOS] In some embodiments, the composition further comprises a second nucleic acid molecule comprising a second nucleic acid sequence encoding a single-chain polypeptide trimer. The single- chain polypeptide trimer encoded by the second nucleic acid sequence comprises a first region, a second region, and a third region. The first region comprises an antigen or antigenic determinant thereof, the second region comprises p2microglobulin or a variant thereof, and the third region comprises a second MHC class I allele or variant thereof. The antigen or antigenic determinant thereof is encoded by an antigen nucleic acid sequence within the second nucleic acid sequence, the P2microglobulin or a variant thereof is encoded by a p2 nucleic acid sequence within the second nucleic acid sequence, and the second M HC class 1 allele or variant thereof is encoded by an MHC nucleic acid sequence within the second nucleic acid sequence. In some embodiments, the antigen or antigenic determinant thereof associates to the second MHC class I allele or variant thereof with an IC50 of no greater than about 500 11M. In some embodiments, the first MHC class I allele is different than the second MHC class I allele.
[0109] In some embodiments, the composition further comprises a third nucleic acid molecule comprising a third nucleic acid sequence encoding a single-chain polypeptide trimer. The singlechain polypeptide trimer encoded by the third nucleic acid sequence comprises a first subunit, a second subunit, and a third subunit. The first subunit comprises an antigen or antigenic determinant thereof, the second subunit comprising P2microglobulin or a variant thereof, and the third subunit comprising a third MHC class I allele or variant thereof. The antigen or antigenic determinant thereof is encoded by an antigen nucleic acid sequence within the third nucleic acid sequence, theDOCKET NO. STFD-OH-PCT PROVISIONAL PATENTP2tnicroglobulin or a variant thereof is encoded by a p2 nucleic acid sequence within the third nucleic acid sequence, and the third MHC class I allele or variant thereof is encoded by an MHC nucleic acid sequence within the third nucleic acid sequence. In some embodiments, the antigen or antigenic determinant thereof associates to the third MHC class I allele or variant thereof with an IC50 of no greater than about 500 nM. In some embodiments, first MHC class I allele is different than the second class I allele, and the third class I allele is different than the first and second class I alleles
[0110] In some embodiments, the composition further comprises a fourth nucleic acid molecule comprising a fourth nucleic acid sequence encoding a single-chain polypeptide trimer. The singlechain polypeptide trimer encoded by the fourth nucleic acid sequence comprises a first subunit, a second subunit, and a third subunit. The first subunit comprises an antigen or antigenic determinant thereof, the second subunit comprises p2microglobulin or a variant thereof, and the third subunit comprises a fourth MHC class I allele or variant thereof. The antigen or antigenic determinant thereof is encoded by an antigen nucleic acid sequence within the fourth nucleic acid sequence, the P2microglobulin or a variant thereof is encoded by a P2 nucleic acid sequence within the fourth nucleic acid sequence, and the fourth MHC class I allele or variant thereof is encoded by an MHC nucleic acid sequence within the fourth nucleic acid sequence. In some embodiments, the antigen or antigenic determinant thereof associates to the fourth MHC class I allele or variant thereof with an IC50 of no greater than about 500 nM. In some embodiments, first MHC class I al lele i s different than the second class I allele, and the third class I allele is different than the first and second class I alleles, and the fourth MHC class I allele is different than the first, second, and third class I alleles
[0111] In some embodiments, the first MHC class I allele is selected from a first family of class I alleles, the second MHC class I allele is from a second family of class I alleles, and the first family of class I alleles is different than the second family of class I alleles. In some embodiments, the first MHC class I allele is selected from a first family of class I alleles, the second MHC class I allele is from a second family of class I alleles, the third MHC class I allele is selected from a third family of class I alleles, the first family of class I alleles is different than the second family of class I alleles, and the third family of class I alleles is different than the first family of class I alleles andDOCKET NO. STFD-OH-PCT PROVISIONAL PATENT the second family of class I alleles. In some embodiments, the first MHC class I allele is selected from a first family of class I alleles, the second MHC class I allele is from a second family of class I alleles, the third MHC class I allele is selected from a third family of class I alleles, and the fourth MHC class I allele is selected from a fourth family of class I alleles, and the first family of class I alleles is different than the second family of class I alleles, the third family of class I alleles is different than the first family of class I alleles and the second family of class I alleles, and the fourth family of class I alleles is different than the first, second, and third families.
[0112] In some embodiments of the composition, the respective antigen or antigenic determinant thereof has an E-score from about 3.2 to about 5.
[0113] A method of calculating E-score is set forth in the below Methods section within the Examples section. Generally, a plurality of nucleic acids encoding a plurality of single-chain polypeptide trimers leads to some of the peptide-MHC combinations to accumulate on the cell surface. Sorting of cells based on cell surface level of pMHC trimer and deep sequencing of the library pool generates the E-score
[0114] In some embodiments of the composition, the first MHC class I allele is chosen from HLA- A alleles, and the second, the third and / or the fourth MHC class I alleles are chosen from HLA-B or HLA-C alleles In some embodiments, the first MHC class [ allele is chosen from HLA-A*01 alleles, and the second MHC class I allele is chosen from HLA-A*02, HLA-A*03, HLA-A*024, HLA-A*026, HLA-B, and HLA-C alleles. In some embodiments, the first MHC class I allele is chosen from HLA-A*01 alleles, the second MHC class I allele is chosen from HLA-A*02, HLA- A*03, HLA-A*024, and HLA- A* 026 alleles, the third MHC class I allele is chosen from HLA-B alleles; and the fourth MHC class I allele is chosen from HLA-C alleles.
[0115] In some embodiments, at least one of the first nucleic acid molecule, the second nucleic acid molecule, the third nucleic acid molecule, or the fourth nucleic acid molecule further comprises a first flexible nucleic acid sequence and second flexible nucleic acid sequence. Together, the first and second flexible nucleic acid sequences encode a flexible linker. In some embodiments, the first flexible nucleic acid sequence is positioned between and adjacent to the respective antigen nucleic acid sequence and the respective p2 nucleic acid sequence In someDOCKET NO. STFD-011-PCT PROVISIONAL PATENT embodiments, the second flexible nucleic acid sequence is positioned between and adjacent to the 3’ end of the respective p2 nucleic acid sequence and the respective MHC nucleic acid sequence.
[0116] In some embodiments, at least one of the first, nucleic acid sequence, the second nucleic acid sequence, the third nucleic acid sequence, and the fourth nucleic acid sequence further comprise a reporter nucleic acid sequence encoding a reporter. In some embodiments, the reporter is GF P.
[0117] In some embodiments, the composition further comprises a third flexible nucleic acid sequence encoding a third flexible linker. In some embodiments, the reporter nucleic acid sequence is positioned 3’ relative to the MHC nucleic acid sequence, and the third flexible nucleic acid sequence is between and the MHC nucleic acid sequence and the reporter nucleic acid sequence In some embodiments, the third flexible nucleic acid sequence joins the MHC nucleic acid sequence and the reporter nucleic acid sequence.
[0118] In some embodiments, at least one of the first nucleic acid sequence, the second nucleic acid sequence, the third nucleic acid sequence, and / or the fourth nucleic acid sequence further comprise a nucleic acid comprising a linker.
[0119] The disclosure relates, in some embodiments, to an expressible nucleic acid sequence comprising, in addition to the first nucleic acid sequence comprising the elements of claim 1 , a nucleic acid sequence encoding a linker domain comprising a linker peptide, wherein the nucleic acid sequence is positioned between the first region and the second region or between the second region and the third region, or between each of the first and second, and the second and third regions in the 5’ to 3" orientation. Any type of linker or linker peptide can be used The term “linker’1or “linker peptide” is used interchangeable herein.
[0120] In some embodiments, each linker or linker peptide is independently selectable from about 0 to about 25, about 1 to about. 25, about 2 to about 25, about 3 to about 25, about 4 to about 25, about 5 to about 25, about 6 to about 25, about 7 to about 25, about 8 to about 25, about 9 to about 25, about 10 to about 25, about I I to about 25, about 12 to about 25, about 13 to about 25, about 14 to about 25, about 15 to about 25, about 16 to about 25, about 17 to about 25, about 18 to about16DOCKET NO. STFD-011-PCT PROVISIONAL PATENT25, about 19 to about 25, about 20 to about 25, about 21 to about 25, about 22 to about 25, about 23 to about 25, about 24 to about 25 natural or non-natural amino acids in length.
[0121] In some embodiments, each linker or linker peptide is about 0, about I, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25 natural or non- natural amino acids in length. In some embodiments, each linker or linker peptide is independently selectable from a linker or linker peptide that is about 0, about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11 , about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25 natural or non-natural amino acids in length. In some embodiments, each linker or linker peptide is about 21 natural or non-natural amino acids in length.
[0122] In some embodiments, the length of each linker or linker peptide is different For example, in some embodiments, the length of a first linker or linker peptide is about 0, about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25 natural or non-natural amino acids in length, and the length of a second linker is about 0, about 1 , about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25 natural or non-natural amino acids in length, where the length of the first linker is different from the length of the second linker. Various configurations can be envisioned by the present disclosure, where the linker domain comprises L 2, 3, 4, 5, 6, 7, 8, 9, 10 or more linkers or linker peptides wherein the linkers or linker peptides are of similar or different lengths
[0123] In some embodiments, the linker is a 2A linker. In some embodiments, first 2A nucleic acid sequence encoding a first 2 A sequence joined to a selection nucleic acid sequence encoding a selection marker. In some embodiments, the selection marker is an antibiotic resistance protein or regulatory sequence. In some embodiments, the selection nucleic acid sequence encodes a puromycin N-acetyl -transferase. In some embodiments, the respective first 2A nucleic acidDOCKET NO. STFD-OH-PCT PROVISIONAL PATENT sequence and respective selection nucleic acid sequence are positioned 3’ from the respective MHC nucleic acid sequence.
[0124] In some embodiments, at least one of the first, nucleic acid sequence, the second nucleic acid sequence, the third nucleic acid sequence, and the fourth nucleic acid sequence further comprise a second 2A nucleic acid sequence encoding a second 2A sequence joined to the reporter nucleic sequence and following the MHC nucleic acid sequence. In some embodiments, the second 2A nucleic acid sequence is joined to the reporter nucleic acid sequence In some embodiments, the second 2A nucleic acid sequence is interposed between the MHC nucleic acid sequence and the first 2A nucleic acid sequence.
[0125] In some embodiments, at least one of the first MHC nucleic acid sequence, the second MHC nucleic acid sequence, the third MHC nucleic acid sequence, and the fourth MHC nucleic acid sequence encodes an HLA class I allele or variant thereof.
[0126] In some embodiments, at least one of the first nucleic acid sequence, the second nucleic acid sequence, the third nucleic acid sequence, and the fourth nucleic acid sequence further comprise a respective vector backbone nucleic acid sequence. In some embodiments, the respective vector backbone nucleic acid sequence is a lentivector sequence.
[0127] In some embodiments, the first MHC class I allele is an HLA class I allele chosen from one or a combination of two or more of: A*01 :01, A*02:01, A*02:05, A*02: 12, A*03 :01, A*11:01, A*23 :01 , A*24:02, A*30:01 , A*31 :01, A*31 : 08, A*34:01 , A*33:03, A*68:01 , B*07:02, B*08:0l , B*08:02, B* 13:01, B*15:01, B* 15:02, B*15:03, B*18:01, B*27:01, B*27:05, B*27:02, B*35:01, B*35:02, B*35:08, B*39:06, B*40:01, B*40:06, B*40: 10, B*41 :01 , 8*44 0, B*46:01, B*48:03, B*50:01, B*51 :01, B*52:01, 6*54:01, B*56:01 , 6*57:01, B*58:0I, C*03:04, C*04:01 , C*06:02, C*07:01, C*07:02, C*08:02, C* 12:0.3 In some embodiments, the second MHC class I allele is an HLA class I allele chosen from one or a combination of two or more of: A*()l :()l, A*02:01, .V 02.05, A*02: 12, A*03:01, A* 11 :01, A*23:01, A*24:02, A*30:01, A*31 :01, A*31 :08, A*34:01, A*33:03, A*68:01 , B*07:02, 8*08:01, B*08:02, B*13:()l, B*15:01, B*15:02, B* 15:03, B*18:01, B*27:01, B*27:05, B*27:02, B*35:01. B*35:02, B*35:08, B*39:06, B*40:01, B*40:06, B*40: 10, B*41:01, B*44:0, B*46:01, B*48:03, B*50:0l, B*51 :01, B*52:0 l , B*54:01 , B*56:01 , B*57:01 ,DOCKET NO. STFD-011-PCT PROVISIONAL PATENTB*58:01, C*03:04, C*04:01, C*06:02, C*07:01, C*07:02, C*08:02, C*12:03. In some embodiments, the third MHC class I allele is an HLA class I allele chosen from one or a combination of two or more of: A*01 :01, A*02:01, A*02:05, A*02:12, A *03:01, A* ll:01, A*23:01, A*24:02, A*30:01, A*31 :01, A*31 :08, A*34:01, A*33:03, A*68:01, B*07:02, B*08:01, 8’ 08.02, B* 13:01, B*15:01, B*15:02, B* 15:03, B* I 8:01 , B*27:01, B *27:05, B*27:02, B*35:01, 8*35.02. B*35:08, B*39:06, 8*40:01, B*40:06, B*40: 10, B*41 :01, B*44:0, B*46:01, 8*48:03. B*50:01, B*51 :01, B*52:01 , B*54:01, B*56:01 , B*57:01, B*58:01 , C*03:04, C*04:01 , C*06:02, C*07:01, C*07:02, C*08:02, C* 12:03 In some embodiments, the fourth MHC class I allele is an HLA class I allele chosen from one or a combination of two or more of: A*01:01, A*02:01, A *02 05, A*02: 12, A*03 :01 , A* 11 :01 , A*23 :01, A*24:02, A*3():()l, A*31 :01, A*31 :08, A*34:01 , A*33:03, A*68:01, B*07:02, B*08:01, B*08:02, B*13:01, B*15:01. B* 15:02, B* 15:03, B*18:01, B*27:01 , B*27:05, B*27:02, B*35:O1, B*35:02, B*35:08, B*39:06, 8*40.01 , B*40:06, 8*40: 10. 8*41 :01, 8*44:0. 8*46:01. B*48:03, 8*50:01 , 8*51 :01, B*52:01, 8*54:01. B*56:01, B*57:01. B*58:01, C*03:04, C*04:01, C*06:02, C*07:01 , C*07:02, C*08:02, C* 12:03.[0128| Some embodiments relate to a composition that comprises a first nucleic acid molecule and a second nucleic acid molecule. The first nucleic acid molecule comprising a first nucleic acid sequence encoding a first single-chain polypeptide trimer comprising a first subunit, a. second subunit, and a third subunit. The first subunit comprises an antigen or antigenic determinant thereof, the second subunit comprises P2microglobulin or a variant thereof, and the third subunit comprises a first MHC class I allele or variant thereof. The antigen or antigenic determinant thereof is encoded by an antigen nucleic acid sequence within the first nucleic acid sequence, the p2microglobulin or a variant thereof is encoded by a 32 nucleic acid sequence within the first nucleic acid sequence, and the first MHC class I allele or variant thereof is encoded by an MHC nucleic acid sequence within the first nucleic acid sequence. The second nucleic acid molecule comprises a second nucleic acid sequence encoding a second single-chain polypeptide trimer comprising a first subunit, a second subunit, and a third subunit. The first subunit comprises an antigen or antigenic determinant thereof, the second subunit comprises p2microglobulin or a variant thereof, and the third subunit comprises a second MHC class I allele or variant thereof TheDOCKET NO. STFD-OH-PCT PROVISIONAL PATENT antigen or antigenic determinant thereof is encoded by an antigen nucleic acid sequence within the second nucleic acid sequence, the |32microglobulm or a variant thereof is encoded by a p2 nucleic acid sequence within the second nucleic acid sequence, and the second MHC class I allele or variant thereof is encoded by an MHC nucleic acid sequence within the second nucleic acid sequence. The first single-chain polypeptide trimer is different than the second single-chain polypeptide trimer.
[0129] In some embodiments, the composition comprises a third nucleic acid molecule comprising a third nucleic acid sequence encoding a third single-chain polypeptide trimer. The third singlechain polypeptide trimer comprises a first subunit, a second subunit, and a third subunit. The first subunit comprises an antigen or antigenic determinant thereof, the second subunit comprises P2microglobulin or a variant thereof, and the third subunit comprises a first MHC class I allele or variant thereof. The antigen or antigenic determinant thereof is encoded by an antigen nucleic acid sequence within the third nucleic acid sequence, the p2microglobulin or a variant thereof is encoded by a p2 nucleic acid sequence within the third nucleic acid sequence, and the third MHC class I allele or variant thereof is encoded by an MHC nucleic acid sequence within the third nucleic acid sequence. The first single-chain polypeptide trimer is different than the second singlechain polypeptide trimer, and the third single-chain polypeptide trimer is different than the first and the second single-chain polypeptide trimers.|0130] In some embodiments, the composition comprises a fourth nucleic acid molecule comprising a fourth nucleic acid sequence encoding a fourth single-chain polypeptide trimer. The fourth single-chain polypeptide trimer comprises a first subunit, a second subunit, and a third subunit. The first subunit comprising an antigen or antigenic determinant thereof, the second subunit comprising p2microglobulin or a variant thereof, and the third subunit comprising a second MHC class I allele or variant thereof. The antigen or antigenic determinant thereof is encoded by an antigen nucleic acid sequence within the fourth nucleic acid sequence, the P2microglobulin or a variant thereof is encoded by a p2 nucleic acid sequence within the fourth nucleic acid sequence, and the fourth MHC class I allele or variant thereof is encoded by an MHC nucleic acid sequence within the fourth nucleic acid sequence. The first single-chain polypeptide trimer is different thanDOCKET NO. STFD-011-PCT PROVISIONAL PATENT the second single-chain polypeptide trimer, the third single-chain polypeptide trimer is different than the first and the second single-chain polypeptide trimers, and the fourth single-chain polypeptide trimer is different than the first, second, and third single-chain polypeptide trimers.
[0131] In some embodiments, at least one of the respective antigen or antigenic determinant thereof has an E-score from about 3.2 to about 5. In some embodiments, each of the respective antigen or antigenic determinant thereof has an E-score from about 3.2 to about 5.
[0132] In some embodiments, the first MHC class I allele is selected from a first family of class I alleles, the second MHC class I allele is from a second family of class I alleles, the third MHC class I allele is selected from a third family of class I alleles, the fourth MHC class I allele is selected from a fourth family of class I alleles, the first family of class I alleles is different than the second family of class I alleles, the third family of class I alleles is different than the first family of class I alleles and the second family of class I alleles, and the fourth family of class I alleles is different that the first, the second, and the third families of class I alleles.
[0133] In some embodiments, the first. MHC class I allele is chosen from HLA-A alleles, and the second and the third MHC class I alleles are chosen from HLA-B or HLA-C alleles. In some embodiments, the first MHC class I allele is chosen from HLA-A*01 alleles, and the second MHC class I allele is chosen from HLA-A*02, HLA-A*03, HLA-A*024, HLA-A*026, HLA-B, and HLA-C alleles In some embodiments, the first MHC class I allele is chosen from HLA-A*01 alleles, the second MHC class I allele is chosen from HLA-A*02, HLA-A*03, HLA-A*024, and HLA-A*026 alleles, the third MHC class I allele is chosen from HLA-B alleles; and the fourth MHC class I allele is chosen from HLA-C alleles.
[0134] In some embodiments, at least one of the first nucleic acid molecule, the second nucleic acid molecule, the third nucleic acid molecule, or the fourth nucleic acid molecule further comprises a first flexible nucleic acid sequence and second flexible nucleic acid sequence that together encode a flexible linker. In some embodiments, the first flexible nucleic acid sequence is positioned between the respective antigen nucleic acid sequence and the respective fJ2 nucleic acid sequence. In some embodiments, the first flexible nucleic acid sequence is positioned adjacent to the respective antigen nucleic acid sequence and the respective 02 nucleic acid sequence. In someDOCKET NO. STFD-OH-PCT PROVISIONAL PATENT embodiments, the second flexible nucleic acid sequence is positioned between the 35end of the respective p2 nucleic acid sequence and the respective MHC nucleic acid sequence. In some embodiments, the second flexible nucleic acid sequence is positioned adjacent to the 3’ end of the respective p2 nucleic acid sequence and the respective MHC nucleic acid sequence.
[0135] In some embodiments, at least one of the first nucleic acid sequence, the second nucleic acid sequence, the third nucleic acid sequence, and the fourth nucleic acid sequence further comprise a reporter nucleic acid sequence encoding a reporter. In some embodiments, at least one of the first nucleic acid sequence, the second nucleic acid sequence, the third nucleic acid sequence, and the fourth nucleic acid sequence further comprise a third flexible nucleic acid sequence encoding a third flexible linker. In some embodiments, the reporter nucleic acid sequence is positioned 3’ relative to the MHC nucleic acid sequence. In some embodiments, the third flexible nucleic acid sequence between the MHC nucleic acid sequence and the reporter nucleic acid sequence. In some embodiments, the third flexible nucleic acid sequence joins the MHC nucleic acid sequence and the reporter nucleic acid sequence.
[0136] In some embodiments, at least one of the first nucleic acid sequence, the second nucleic acid sequence, the third nucleic acid sequence, and the fourth nucleic acid sequence further comprise a first 2 A nucleic acid sequence encoding a first 2A sequence joined to a selection nucleic acid sequence. In some embodiments, the selection nucleic acid sequence encodes an antibiotic resistance protein or regulatory sequence. In some embodiments, the selection nucleic acid sequence encodes a puromycin N-acetyl-transferase. In some embodiments, the respective first 2A nucleic acid sequence and respective selection nucleic acid sequence are positioned 3’ from the respective MHC nucleic acid sequence.
[0137] In some embodiments, at least one of the first nucleic acid sequence, the second nucleic acid sequence, the third nucleic acid sequence, and the fourth nucleic acid sequence further comprise a second 2A nucleic acid sequence encoding a second 2A sequence joined to the reporter nucleic sequence and following the MHC nucleic acid sequence. In some embodiments, the second 2A nucleic acid sequence joined to the reporter nucleic acid sequence is interposed between theDOCKET NO. STFD-011-PCT PROVISIONAL PATENTMHC nucleic acid sequence and the first 2A nucleic acid sequence. In some embodiments, the reporter nucleic acid sequence encodes a GFP.
[0138] In some embodiments, at least one of the first MHC nucleic acid sequence, the second MHC nucleic acid sequence, the third MHC nucleic acid sequence, and the fourth MHC nucleic acid sequence encodes an Hl .A class I allele or variant thereof
[0139] In some embodiments, at least one of the first nucleic acid molecule, the second nucleic acid molecule, the third nucleic acid molecule, and / or the fourth nucleic acid molecule further comprise a respective vector backbone nucleic acid sequence. In some embodiments, the respective vector backbone nucleic acid sequence is a lentivector sequence. In some embodiments
[0140] In some embodiments, the first MHC class I allele is an HLA class I allele chosen from one or a combination of two or more of: A*01 :01, A*02:01, A*02:05, A*02: 12, A*03:01, A* 11 :01, A*23:01, A*24:02, A*30:01, A*31 :01, A*31 :08, A*34:01, A*33:03, A*68:01, B*07:02, B*08:0L B*08:02, B*13:01, B*15:01, B*15:02, 8*15:03, B*18:01, B*27:01, B*27:05, B*27:02, B*35:01, B*35:02, B*35:08, B*39:06, B*40:01, B*40:06, B*40: 10, B*41:0l, B*44:0, B*46:01 , B*48:03, B*50:01, B*51:01, B*52:01, B*54:01, B*56:01, B*57:01, B*58:01, C*03:04, C*04:01, C*06:02, C*07:01, C*07:02, C*08:02, C*12:03. In some embodiments, the second MHC class I allele is an HLA class 1 allele chosen from one or a combination of two or more of: A*01 :01 , A*02:01 , A*02 :05, A*02: 12, A*03:01, A*ll:01, A*23:01, A*24:02, A*30:01, A*31 :01, A*31 ;08, A*34:01, A* 33 : 03 , A * 68 : 01 , B * 07 : 02, B * 08 : 01 , B * 08 : 02, 8 * 13 : 01 , B * 15 : 01 , B * 15 : 02, B * 15 : 03 , B * 18 : 01 , B*27:01, 8*27:05, B*27:02, B*35:O1, B*35:02, 6*35:08, B*39:06, B*40:01, B*40:06, B*40: 10, 8*41 :01, B *44:0, B*46:01, B*48:03, B*50:01, 8*51 :01, B*52:01 , B*54:01, B*56:01, B*57:01, B*58:01, C*03:04, C*04:01, C*06:02, C*07:01, C*07:02, C*08:02, C*12:03. In some embodiments, the third MHC class I allele is an HLA class I allele chosen from one or a combination of two or more of. A*()l :01, A*02:01, A*02:05, A*02: 12, A*03:01, A* ll :01 , A*23:01, A*24:02, A*30:01, A*31 :01, A*31 :08, A*34:01, A*33:03, A*68:01, 8*07:02, B*08:01, 8*08:02, 13*13:01, 8*15:01 , 13*15:02, 8*15:03, 13*18:01, 8*27:01, 13*27:05, 8*27:02, 13*35:01, B*35:02, B*35:08, B*39:06, B*40:01. B*40:06, B*40: 10, B*41:01, B*44:0, B*46:01, B*48:03, 13*50:01, 8*51:01 , 13*52:01, 8*54:01 , 13*56:01, 8*57:01 , 13*58:01, C*03:04, C*04:01, ('*06:02.DOCKET NO. STFD-011-PCT PROVISIONAL PATENTC*07:01, C*07:02, C*08:02, C*12:03. In some embodiments, the fourth MHC class I allele is an HLA class I allele chosen from one or a combination of two or more of: A*01:01, A*02:01, A*02:05, A*02: 12, A*03:01, A* 11 :0L A*23:01 , A*24:02, A*30:01 , A*31 :0L A*31 :08, A*34:01, A*33:03, A*68:01, B*07:02, B*08:01, B*08:02, B*13:01, B*15:01, B*15:02, B*15:O3, B*18:01, B*27:01, B*27:05, B*27:02, B*35 :01 , B*35:02, B*35:08, B*39:06, B *40:01, B*40:06, B*40: 10, B*41 :01, B*44:0, B*46:01, B*48:03, B*50:01, B*51 :01, B*52:01, B*54:01, B*56:01, B*57:01 , B*58:01, C*03:04, C*04:01, C*06:02, C*07:01, C*07:02, C*08:02, C*12:03.
[0141] Some embodiments comprise any of the above-described compositions where at least one of the first nucleic acid molecule, the second nucleic acid molecule, the third nucleic acid molecule, and the fourth nucleic acid molecule further comprise a signal nucleic acid sequence encoding a signal peptide. In some embodiments, the respective nucleic acid molecule comprises in 5’ to 3’ orientation the signal nucleic acid sequence, the antigen nucleic acid sequence, the p2 nucleic acid sequence, the MHC nucleic acid sequence.
[0142] The disclosure relates to a composition including one of the above-described compositions but where one or more, or all of, the nucleic acid sequences that encode antigens are replaced by a respective insertion site. The disclosure also relates to a composition including one of the abovedescribed compositions but where one or more, or all of, the nucleic acid sequences that encode antigens or antigenic determinants thereof are replaced by one or a plurality of respective insertion sites, each insertion site comprising one or a plurality of restriction enzyme recognition sequences. The disclosure also relates to a composition including one of the above-described compositions but where one or more, or all of, the nucleic acid sequences that encode antigens or antigenic determinants thereof and the nucleic acid sequences that encode the MHC Class I allele or a variant thereof are replaced by one or a plurality of respective insertion sites.
[0143] The disclosure relates io a cell comprising or a composition comprising a nucleic acid molecule that is a plasmid comprising a vector backbone and at least a first multiple cloning site that comprises expressible nucleic acid sequence operably linked to at least one promoter positioned on the vector backbone; wherein the expressible nucleic acid sequence encodes a p2 microglobulin or a variant thereof and a first MHC class I allele or variant thereof In someDOCKET NO. STFD-011-PCT PROVISIONAL PATENT embodiments, the first multiple cloning site comprises an insertion site that, when subcloned with a nucleic acid sequence encoding an antigen or antigenic determinant thereof’ positions the nucleic acid encoding the antigen or antigenic determinant thereof in frame with the nucleic acid sequence encoding a MHC class I allele or variant thereof and a nucleic acid sequence encoding a [52 microglobulin or a variant thereof. In some embodiments, plasmid is a DNA, RNA or DNA / RNA hybrid molecule comprising a promoter region operably linked to the expressible nucleic acid sequence, such that upon exposure to an RNA polymerase in a cell, the cell expresses the polypeptide encoded by the first expressible nucleic acid sequence. In some embodiments, the polypeptide is a trimer comprising: (i) from about 8 to about 20 aminos acids that are the antigen or antigen determinant thereof, (ii) from about 8 to about 100 amino acids that are the MHC class I allele or variant thereof; and (iii) from about 10 to about 200 amino acids that are the [32 microglobulin or a variant thereof.
[0144] Embodiments of the disclosure include one or a collection of two or more of the abovedescribed nucleic acid molecules. Some embodiments comprise one or a plurality of any nucleic acid molecules disclosed herein. Some embodiments comprise one or a collection of two or more of amino acid sequences encoded by the above-described nucleic acid molecules. Some embodiments comprise a one or a collection of any amino acid sequence encoded by a nucleic acid molecule herein.[0145 | C ompositions of the disclosure include a cell or plurality of cells comprising one or more of the above-described compositions. Some embodiments comprise a cell that comprises one or more of any composition herein.
[0146] The disclosure also relates to a kit comprising one or a plurality of nucleic acid molecules disclosed herein In some embodiments, the kit comprises one or more of the above-described compositions. In some embodiments, the kit comprises one or a plurality of cells. In some embodiments, the kit comprises a first container comprising one or a plurality of cells and a second container comprising at least one nucleic acid molecule disclosed herein. In some embodiments, the kit comprises (i) a first container comprising the first nucleic acid molecule disclosed herein, a second container comprising the second nucleic acid molecule disclosed herein, or (ii) a firstDOCKET NO. STFD-011-PCT PROVISIONAL PATENT container comprising the first nucleic acid molecule disclosed herein, a second container comprising the second nucleic acid molecule disclosed herein; and a third container comprising the third nucleic acid molecule disclosed herein; or (iii) a first container comprising the first nucleic acid molecule disclosed herein, a second container comprising the second nucleic acid molecule disclosed herein; a third container comprising the third nucleic acid molecule disclosed herein; and a fourth container comprising the fourth nucleic acid molecule disclosed herein. In some embodiments, the kit comprises (i) a first container comprising one or a plurality of cells and a second container comprising the first nucleic acid molecule disclosed herein, a third container comprising the second nucleic acid molecule disclosed herein; or (ii) a second container comprising the first nucleic acid molecule disclosed herein, a third container comprising the second nucleic acid molecule disclosed herein; and a fourth container comprising the third nucleic acid molecule disclosed herein; or (iii) a second container comprising the first nucleic acid molecule disclosed herein, a third container comprising the second nucleic acid molecule disclosed herein, a fourth container comprising the third nucleic acid molecule disclosed herein; and a fifth container comprising the fourth nucleic acid molecule disclosed herein.
[0147] Some embodiments of a kit or method herein further comprise a staining reagent or use thereof In some embodiments, the staining reagent is a dye conjugated anti-p2-microglobulin antibody. In some embodiments the dye conjugated anti-P2-microglobulin antibody is a PE-anti human p2-microglobulin antibody or a clone 2M2. In some embodiments, the staining reagent is a dye conjugated anti-human HLA antibody. In some embodiments, the anti-human HLA antibody is an anti-IILA-A2 antibody. In some embodiments, the anti-human HLA antibody is limited to certain HLA alleles instead of staining across all of HLA-A, B and C.
[0148] In some embodiments, the cell is capable of being transformed with one or more of the first nucleic acid molecule, the second nucleic acid molecule, the third nucleic acid molecule, and the fourth nucleic acid molecule as described herein. In some embodiments, the cell comprises one or more of the first nucleic acid molecule, the second nucleic acid molecule, the third nucleic acid molecule, and the fourth nucleic acid molecule as described herein. In some embodiments, the cell is a dual HLA and TAP knock-out cell.17DOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0149] In some embodiments, kits in accordance with the present disclosure may be used to culture and / or to propagate cells or cell types of interest. In some embodiments, kits for culturing cells comprise the substrate described above and optionally further comprise medium and a cell type of interest. .Any array, system, or component thereof disclosed may be arranged in a kit either individually or in combination with any other array; system, or component thereof. The invention provides a kit to perform any of the methods described herein. In some embodiments, the kit comprises at least one container comprising one or a plurality of polypeptides comprising a polypeptide sequence associated with the extracellular matrix or functional fragments thereof. In some embodiments, the kit comprises at least one container comprising any of the polypeptides or functional fragments described herein. In some embodiments, the polypeptides are in solution (such as a buffer with adequate pH and / or other necessary' additive to minimize degradation of the polypeptides during prolonged storage). In some embodiments, the polypeptides are lyophilized for the purposes of resuspension after prolonged storage. In some embodiments, the kit comprises: at least, one container comprising one or a plurality of polypeptides comprising a polypeptide sequence associated with the extracellular matrix (or functional fragments thereof); and a solid support upon which the polypeptides or fragments may be affixed. In some embodiments, the kit optionally comprises instructions to perform any or all steps of any method described herein. In some embodiments, the kit comprises an array or system described herein and instructions for implementing one or a plurality of steps using a computer program product disclosed herein. It is understood that one or a plurality' of the steps from any of the methods described herein can be performed by accessing a computer program product encoded on computer storage medium directly through one or more computer processors or remotely through one or more computer processors via an internet connection or other virtual connection to the one or more computer processors. In some embodiments, the kit comprises a computer-program product, described herein or requisite information to access a computer processor comprising the computer program product encoded on computer storage medium remotely. In some embodiments, the computer program product, when executed by a user, calculates one or more adhesion values, normalizes the one or more adhesion values, generates one or more adhesion signatures or one or more adhesion profiles,DOCKET NO. STFD-011-PCT PROVISIONAL PATENT and / or displays any of the adhesion values, adhesion signatures, adhesion profiles to a user. In some embodiments, the kit comprises a computer program product encoded on a computer- readable storage medium that comprises instructions for performing any of the steps of the methods described herein. In some embodiments, the invention relates to a kit comprising instructions for providing one or more adhesion values, one or more normalized adhesion values, one or more adhesion profiles, one or more adhesion signatures, or any combination thereof. In some embodiments, the kit comprises a computer program product encoded on a computer storage medium that when, executed on one or a plurality of computer processors, quantifies an adhesion value, determines an adhesion signature or adhesion profile, and / or displays an adhesion signature, adhesion value, adhesion signature, and / or any combination thereof. In some embodiments, the kit comprises a computer program product encoded on a computer storage medium that, when executed by one or a plurality of computer processors, quantifies adhesion values of one or more cells samples and determines an adhesion signature based at least partially upon the adhesion values. In some embodiments, kit comprises instructions for accessing the computer storage medium, quantifying adhesion values, normalizing adhesion values, determining an adhesion signature of a cell type, and / or any combination of steps thereof. In some embodiments, the computer-readable storage medium comprises instructions for performing any of the methods described herein. In some embodiments, the kit comprises an array or system disclosed herein and a computer program product encoded on computer storage medium that, when executed, performs any of the method steps disclosed herein individually or in combination and provides instructions for performing any of the same steps. In some embodiments, the instructions comprise an instruction to adhere any one or plurality of polypeptides disclosed herein to a solid support.
[0150] The disclosure further provides for a kit comprising one or a plurality of containers that comprise one or a plurality of the polypeptides or fragments disclosed herein. In some embodiments, the kit comprises cell media free of serum, or any animal-based derivative of serum that enhances the culture or proliferation of cells. In some embodiments, the kit comprises: an array disclosed herein, any cell media disclosed herein, and a computer program product disclosed herein optionally comprising instructions to perform any one or more steps of any methodDOCKET NO. STFD-011-PCT PROVISIONAL PATENT disclosed herein. In some embodiments, the kit does not comprise cell media In some embodiments, the kit comprises a solid support free of any one individual pair of polypeptides disclosed herein In some embodiments, the kit comprises a device for affixing one or more adhesion sets to a solid support.
[0151] The kit may contain two or more containers, packs, or dispensers together with instructions for preparation of an array. In some embodiments, the kit comprises at least one container comprising the array or system described herein and a second container comprising a means for maintenance, use, and / or storage of the array such as storage buffer. In some embodiments, the kit comprises a composition comprising any polypeptide disclosed herein or the nucleic acid molecules in solution or lyophilized or dried and accompanied by a rehydration mixture. In some embodiments, the polypeptides and rehydration mixture may be in one or more additional containers.
[0152] T he compositions included in the kit may be supplied in containers of any sort such that the shelf-life of the different components are preserved, and are not adsorbed or altered by the materials of the container. For example, suitable containers include simple bottles that may be fabricated from glass, organic polymers, such as polycarbonate, polystyrene, polypropylene, polyethylene, ceramic, metal or any other material typically employed to hold reagents or food; envelopes, that may consist of foil-lined interiors, such as aluminum or an alloy. Other containers include test tubes, vials, flasks, and syringes The containers may have two compartments that are separated by a readily removable membrane that upon removal permits the components of the compositions to mix. Removable membranes may be glass, plastic, rubber, or other inert material
[0153] Kits may also be supplied with instructional materials. Instructions may be printed on paper or other substrates, and / or may be supplied as an electronic-readable medium, such as a floppy disc, CD-ROM, DVD-ROM, flash drive, zip disc, videotape, audio tape, or other readable memory storage device. Detailed instructions may not be physically associated with the kit; instead, a user may be directed to an internet web site specified by the manufacturer or distributor of the kit, or supplied as electronic mail.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0154] The disclosure also provides a kit comprising: an array of polypeptides, the array comprising: a solid support and a plurality of trimers disclosed herein, wherein each trimer comprises a MHC cl ass I alleles or variant thereof; and optionally comprising a cell culture vessel . In some embodiments, the kit further comprises at least one of the following: cell media, a volume of fluorescent stain or dye, a cell sample, and a set of instructions, optionally accessible remotely through an electronic medium.Methods
[0155] Some embodiments comprise a method of identifying an immunotherapy target. The method comprises selecting from a population of cells comprising one or more of the compositions described herein a subpopulation of cells surface displaying the respective single-chain polypeptide trimer. The method also comprises identifying the respective antigen or antigenic determinant thereof within the respective single-chain polypeptide trimer for at least one cell in the subpopulation. In some embodiments, the step of identifying comprises identifying the respective antigen or antigenic determinant thereof within the respective single-chain polypeptide trimer for each cell in the subpopulation by exposing a cell expressing the disclosed trimer or trimers with an antibody specific for an amino acid sequence on the trimer or a stain specific for an amino acid sequence within the primer positioned on or within the surface cell membrane or lipid bilayer. In some embodiments, the methods of identifying an immunotherapeutic protein or epitope comprise:(a) transfecting a cell or plurality7of cells with one or a plurality7of nucleic acid molecules disclosed here,(b) culturing the cell or plurality of cells for a time period sufficient to allow7expression of the single-chain trimer;(c) detecting and / or quantifying the expression of the trimer on the surface of the cell or plurality of cells.In some embodiments, the method further comprises a step of calculating an E-score for each identified antigen or antigenic determinant. In some embodiments, the method further comprises selecting an identified antigen or antigenic determinant thereof as a therapeuticDOCKET NO. STFD-011-PCT PROVISIONAL PATENT target or epitope if the calculated E-score is a high E-score. In some embodiments, the high E-score is about 3.2 to about 5.
[0156] In some embodiments an E-score is calculated as follows. To calculate E-score, first all the reads are normalized within each bin through dividing the counts in each bin by the average to account for sequencing read depth. Subsequently, the total number of sequencing reads measured per trimer (pHLA pair) were normalized. This step is to adjust for uneven distribution of antigen- HLA reads that were originated from various steps in the ESCAPE-seq experiments. To synthesize these data into a singular E-score, the following formula was used: Escore::::counts bg * w bg ■r counts low *w low t- counts med * w med+counts high* w high where w bg:::0,iv low = 2,w jned =4,w high::::8, where weight was put at log scale that matched with the binning scale during cell sorting, while assigning background bin as 0.
[0157] In some embodiments an identified antigen or antigenic determinant is the immunotherapy target. In some embodiments, the identifying comprises sequencing the antigen nucleic acid sequence.
[0158] In some embodiments, the method further comprises transfecting cells with any one or more of the above-described compositions. In some embodiments, the method further comprises transfecting cells a nucleic acid encoding any single-chain trimer herein
[0159] In some embodiments, the method further comprises cloning an antigen nucleic acid sequence into the insertion site of a composition herein lacking the antigen nucleic acid sequence but comprising an insertion site.
[0160] In some embodiments, the method further comprises transfecting cells with a nucleic acid molecule library, wherein a plurality of nucleic acid molecules in the nucleic acid library each comprise a nucleic acid sequence encoding a single-chain polypeptide trimer. In some embodiments, the single-chain polypeptide trimer comprises a first subunit, a second subunit, and a third subunit. The first subunit comprises an antigen or antigenic determinant thereof, the second subunit comprises P2microglobulin or a variant thereof, and the third subunit comprises an MIIC class I allele or variant thereof. The antigen or antigenic determinant thereof is encoded by an antigen nucleic acid sequence within the nucleic acid sequence, the p2microglobulin or a variantDOCKET NO. STFD-011-PCT PROVISIONAL PATENT thereof is encoded by a p2 nucleic acid sequence within the nucleic acid sequence, and the first MHC class I allele or variant thereof is encoded by an MHC nucleic acid sequence within the nucleic acid sequence. Each of the antigen nucleic acid sequence, the p2 nucleic acid sequence, and the MHC nucleic acid sequence comprises a respective 5’ end and respective 3’ end.
[0161] In some embodiments, two or more nucleic acid molecules in the nucleic acid library differ from each other in at least one of the antigen nucleic acid sequences and the MHC nucleic acid sequence. In some embodiments, the MHC nucleic acid sequence in different ones of the two or more nucleic acid molecules encodes an HLA-A, HLA-B, or HLA-C allele. In some embodiments, the HLA-A, HLA-B, or HLA-C allele are independently and respectively selected from A*01:01, A*02:01, A*02:05, A*02:12, A*03:01, A*ll:01, A*23:01, A*24:02, A*30:01, A*31:01, A*31 .08, A*34:01, A*33:03, A*68:01, B*07:02, B*08:01, B*08:02, B*13:01, B* 15:01, B* 15:02, B*15:03, B*18:01 , B*27:01, B*27:05, B*27:02, 8*35:01 , B*35:02, B*35:08, B*39:06, B*40:01 , B*40:06, 8*40:10, B*41 :01, 8*44:0, B*46:01, B*48:03, 8*50:01, B*51 :01, B*52:01, B*54:01, B*56:01, B*57:01, B*58:01, C*03:04, C*04:01, C*06:02, C*07:01, C*07:02, C*08:02, or C*12:03. In some embodiments, a set of the different ones of the two or more nucleic acid molecules encodes each of A*01 :01, A*02:01, A*02:05, A*02:12, A*03:01, A* 11:01, A*23:01, A*24:02, A*30:01, A*31 :01 , A*31 :08, A*34:01, A*33:03, A*68:01, B*07:02, 8*08:01 , B*08:02, 8*13:01 , B* 15:01 , B* 15:02, B*15:03, B*18:01, B*27:01, B*27:05, B*27:02, B*35:O1, B*35:02, B*35:08, B*39:06, B*40:01, B*40:06, 8*40: 10, 8*41 :01, B*44:0, B*46:01, B*48:03, 8*50:01, 8*51 :01, B*52:01, B*54:01, 8*56:01, B*57:01, 8*58:01, C*03:04, C*04:01, C*06:02, C*07:01, C*07:02, C*08:02, and C* 12:03 on separate ones of the two or more nucleic acid molecules. In some embodiments, a set of the different ones of the two or more nucleic acid molecules encodes different ones of two or more of an HLA-A allele, an HLA-B allele, or HLA-C allele.
[0162] In some embodiments of the method, the nucleic acid molecule further comprises a first flexible nucleic acid sequence and second flexible nucleic acid sequence that together encode a flexible linker. In some embodiments, the first flexible nucleic acid sequence is positioned between the antigen nucleic acid sequence and the P2 nucleic acid sequence. In some embodiments, the first flexible nucleic acid sequence is positioned adjacent to the antigen nucleic acid sequence andDOCKET NO. STFD-011-PCT PROVISIONAL PATENT the p2 nucleic acid sequence. In some embodiments, the second flexible nucleic acid sequence is positioned between the 3 ’ end of the £>2 nucleic acid sequence and the MHC nucleic acid sequence. In some embodiments, the second flexible nucleic acid sequence is positioned adjacent to the 3’ end of the P2 nucleic acid sequence and the MHC nucleic acid sequence. In some embodiments, the nucleic acid sequence further comprises a reporter nucleic acid sequence encoding a reporter. In some embodiments, the nucleic acid molecule further comprises a third flexible nucleic acid sequence encoding a third flexible linker. In some embodiments, the reporter nucleic acid sequence is positioned 3 ’ relative to the MHC nucleic acid sequence. In some embodiments, the third flexible nucleic acid sequence is between the MHC nucleic acid sequence and the reporter nucleic acid sequence. In some embodiments, the third flexible nucleic acid sequence joins the MHC nucleic acid sequence and the reporter nucleic acid sequence. In some embodiments, the nucleic acid sequence further comprises a first 2A nucleic acid sequence encoding a first 2A sequence joined to a selection nucleic acid sequence. In some embodiments, the selection nucleic acid sequence encodes an antibiotic resistance protein or regulatory sequence. In some embodiments, the selection nucleic acid sequence encodes a puromycin N-acetyl -transferase. In some embodiments, the first 2A nucleic acid sequence and selection nucleic acid sequence are positioned 3’ from the MHC nucleic acid sequence. In some embodiments, the nucleic acid sequence further comprises a second 2A nucleic acid sequence encoding a second 2A sequence joined to the reporter nucleic sequence and following the MHC nucleic acid sequence. In some embodiments, the second 2A nucleic acid sequence joined to the reporter nucleic acid sequence is interposed between the MHC nucleic acid sequence and the first 2A nucleic acid sequence. In some embodiments, the reporter nucleic acid sequence encodes a GFP. In some embodiments, the MHC nucleic acid sequence encodes an HL A class I allele or variant thereof. In some embodiments, the nucleic acid sequence further comprises a vector backbone nucleic acid sequence. In some embodiments, the vector backbone nucleic acid sequence is a lentivector sequence.
[0163] In some embodiments, the step of selecting comprises detecting cells with surface expressed p2microglobulin and / or MHC class I allele.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0164] In some embodiments, the method further comprises inserting a plurality of antigen nucleic acid sequences into an insertion site of a nucleic acid cassette comprising the P2 nucleic acid sequence within the nucleic acid sequence and the MHC nucleic aci d sequence to create the nucleic acid molecule.
[0165] In some embodiments, the antigen nucleic acid sequence is selected from coding sequences of one or more of a viral protein, an oncoprotein, or a non-viral intracellular pathogen protein.
[0166] In some embodiments, the cells are HLA / TAP knockout cells In some embodiments, the antigen or antigenic determinant thereof is 8-10 amino acids in length.
[0167] The disclosure also relates to a method of predicting or quantifying affinities of a panel of antigen sequences associated with an epitope or an HLA molecule comprising exposing any singlechain trimer disclosed herein comprising an antigen sequence with an unknown binding affinity to an epitope or HLA molecule with the epitope sequence or the HLA molecule; calculating an E- score as disclosed herein, and repeating step (i) and (ii) in respect to other members of the panel of antigens with their corresponding single-chain trimers disclosed herein comprising the other members of the panel, compiling the relative affinities of each antigen sequence to the epitope or the HLA molecule; and ranking the affinities relative to each other sequences in the panel, such that higher E-scores are indicative of more tightly bound protein-protein interactions and lower E- scores are indicative of less tightly bound protein-protein interactions between the antigen or amino acid sequence and the epitope or HLA molecule. In some embodiments, the method further comprises predicting the binding affinity of an amino acid sequence relative to an epitope or HLA molecule by characterizing the binding affinity of each amino acid sequence within the panel based upon the ranking. In some embodiments, the step of calculating the E-score is followed by normalizing the E-scores assigned to each amino acid in the relative to a control affinity value (taken from a known binder or known non-binder to the epitope or HLA molecule) and then performing the step of ranking the amino acid sequences in order of quantitative binding affinities to the epitopes of HLA molecules to which they were exposed In some embodiments, the method of predicting quantitative binding affinity of an amino acid sequence with an epitope or HLA molecule, wherein the step of exposing is followed by calculating an E-score; which is followedDOCKET NO. STFD-011-PCT PROVISIONAL PATENT by a step of characterizing the E-score with an affinity relative to an E-score of a known binding interaction between a control amino acid sequence and the epitope or HLA molecule, such that higher E-scores than the control amino acid are more predictive of a tightly bound protein -protein interaction between the HLA molecule and the amino acid sequence or antigen sequence, and an E-score corresponding to the amino acid or antigen lower than the E-score of the control is predictive of a less tightly bound protein-protein interaction between the amino acid or antigen and the epitope or HLA molecule, relative to the control binding affinity. In some embodiments, the method further comprises transfecting cells with any one or more of the above-described compositions comprising a nucleic acid encoding the single-chain trinier comprising the amino acid sequence. In some embodiments, the method further comprises transfecting cells with a nucleic acid encoding a single-chain trimer comprising a control amino acid with a known affinity to the epitope.
[0168] T he disclosure relates to a method of predicting quantitative binding affinities among a first and second amino acid sequence comprising: (i) exposing the first amino acid sequence positioned within a disclosed single-chain trimer to an epitope; (ii) exposing the second amino acid sequence positioned within a disclosed single-chain trimer to the same epitope; (iii) calculating the E-score of the first amino acid and the second amino acid relative to the epitope; comparing the E- score of the first amino acid with the E-score of the second amino acid and predicting the quantitative affinity of the first amino acid to the epitope relative to the second amino acid, such that if the first E-score is higher than the second E-score, the first amino acid is a stronger binder to the epitope than the second amino acid. In some embodiments, the first amino acid sequence is from about 8 to about 24 amino acids in length as a component of the single-chain trimer. In some embodiments, the epitope is a viral epitope or a tumor associated antigen.
[0169] The disclosure also relates to a method of predicting quantitative binding affinities among a first and second amino acid sequence comprising: (i) exposing the first amino acid sequence positioned within a disclosed single-chain trimer to an HLA molecule; (ii) exposing the second amino acid sequence positioned within a disclosed single-chain trimer to the same HLA molecule; (hi) calculating the E-score of the first amino acid and the second amino acid relative to the HLADOCKET NO. STFD-011-PCT PROVISIONAL PATENT molecule; comparing the E-score of the first amino acid with the E-score of the second amino acid and predicting the quantitative binding affinity of the first amino acid to the HLA molecule relative to the second amino acid, such that if the first E-score is higher than the second E-score, the first amino acid is a stronger binder to the HLA molecule than the second amino acid. In some embodiments, the first amino acid sequence is from about 8 to about 24 amino acids in length as a component of the single-chain trimer. In some embodiments, the HLA molecule is a human HLA molecule disclosed herein. In some embodiments, the method further comprises a step of correlating the E-score of the first amino acid to the ability or probability that the first amnio acid is presented on the surface or secreted out of the surface of a cell by the HL A molecule if both the first amino acid sequence and the HLA molecule are co-expressed in a cell.
[0170] The disclosure relates to a method of predicting quantitative binding affinity of a first amino acid sequence to a HL, A molecule in a cell comprising: (i) exposing the first amino acid sequence positioned within a disclosed single-chain trimer disclosed herein to an HLA molecule; and (ii) calculating an E-score of the first amino acid In some embodiments, the method comprises a step of comparing the E-score of the first amino acid to the E-score of an amino acid with a known binding affinity to the HLA molecule and predicting the quantitative binding affinity of the first amino acid to the HLA molecule relative to the E-score of the amino acid with the known binding affinity, such that if the first E-score is higher than the known E-score, the first amino acid is a stronger binder to the HLA molecule than the amino acid with the known binding affinity to the HLA molecule. In some embodiments, the first amino acid sequence is from about 8 to about 24 amino acids in length as a component of the single-chain trimer disclosed herein. In some embodiments, the HL.A molecule is a human HLA molecule disclosed herein. In some embodiments, the method further comprises a step of correlating the E-score of the first amino acid to the ability or probability that the first amnio acid is presented on the surface or secreted out of the surface of a cell with the HL.A molecule if both the first amino acid sequence and the HLA molecule are co-expressed in the cell. In some embodiments, the step of exposing is performed simultaneously with a panel of HLA peptides and a panel of other amino acid sequences, wherein each of the other amin acid sequences are positioned within a respective, independently selectable18DOCKET NO. STFD-011-PCT PROVISIONAL PATENT single-chain turner disclosed herein In some embodiments, the number of other amino acid sequences is no less than about 100, 500, 1,000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, or 80,000 or more other amino acids. In some embodiments, the method further comprises exposing the first amino acid to a first HLA molecule in a first well of a multiwell system and exposing the first amino acid to a second HLA molecule in a second well of the multiwell system [0171} Some embodiments comprise a system. The system comprises one or a plurality of one or more composition herein, or one or more nucleic acid herein, and HLA / TAP knockout cells. . In some embodiments, the method further comprises exposing the first amino acid to a first HLA molecule in a first well of a multiwell system and, concurrently, exposing the first amino acid to a panel of HLA molecules, each HLA molecule paired with the first amino acid sequence in its own well of the multiwell system. In some embodiments, the multiwell system comprises at least one cell, cand each cell comprises a nucleic acid molecule comprising a nucleic acid sequence that encodes a single chain trimer disclosed herein comprising the first amino acid sequence and one HLA molecule In some embodiments, the first amino acid sequence or plurality of amino acid sequences are viral peptide, tumor associated antigens or variants thereof.
[0172] Some embodiments comprise a method of making the composition herein. The method comprises cloning one or more antigen sequences into the insertion site of the composition free of an antigen sequence. In some embodiments, the one or more antigen sequence comprises a library' of antigen sequences.
[0173] Some embodiments comprise a composition comprising any one or more antigen or antigenic determinant thereof disclosed herein In some embodiments, the composition further comprises a match of any one or more antigen or antigenic determinant thereof to one or more HLA molecules pairwise in a well, such that the composition comprises a solid support such as a plastic, having a plurality of independent addressable wells, and each well comprising a singlechain trimer disclosed herein comprising a single antigen, amino acid or antigenic determinant thereof paired with a single HLA molecule. In some embodiments, there are no less than about 96 wells, 1,000 wells or a systems with multiple solid supports, each solid support comprising a plurality of wells, such that the system comprises a total number of well of about 10,000, 20,000,DOCKET NO. STFD-011-PCT PROVISIONAL PATENT30,000, 40,000, 50,000, 60,000, 70,000, 75,000 or more wells. In some embodiments, each well comprises a cell comprising a first nucleic acid sequence encoding the antigen, antigenic determinant thereof or amino acid disclosed herein and a second nucleic acid sequence encoding the epitope or HLA molecule. In some embodiments, methods of the disclosure are performed using the system such that many thousands of E-scores are calculated simultaneously or serially, and a ranking of those E-scores are assigned to each amino acid, antigen, or antigenic determinant thereof.
[0174] Some embodiments comprise a pharmaceutical composition comprising (1) any antigen or antigenic determinant thereof disclosed herein and (2) a pharmaceutically acceptable carrier.
[0175] Some embodiments comprise a dual HLA and TAP knock-out cell. Some embodiments comprise a composition comprising a dual HLA and TAP knock-out cell. In some embodiments, the cell comprises any one or more nucleic acid disclosed herein. In some embodiments, the cell comprises any one or more amino acid sequence disclosed herein.
[0176] Some embodiments comprise a composition comprising any one or more amino acid sequence disclosed herein directly or translated from nucleic acid sequence.
[0177] Some embodiments comprise a pharmaceutical composition comprising any one or more amino acid sequence disclosed herein directly or translated from nucleic acid sequence and a pharmaceutically acceptable carrier
[0178] Although the description of pharmaceutical compositions provided herein are principally directed to pharmaceutical compositions which are suitable for ethical administration to humans, it will be understood by the skilled arts san that such compositions are generally suitable for administration to animals of all sorts. Modification of pharmaceutical compositions suitable for administration to humans in order to render the compositions suitable for administration to various animals is well understood, and the ordinarily skilled veterinary pharmacologist can design and perform such modification with merely ordinary, if any, experimentation. Subjects to which administration of the pharmaceutical compositions of the invention is contemplated include, but are not limited to, humans and other primates, mammals including commercially relevant mammals such as non-human primates, cattle, pigs, horses, sheep, cats, and dogs.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0179] Pharmaceutical compositions herein may be prepared, packaged, or sold in formulations suitable for ophthalmic, oral, rectal, vaginal, parenteral, topical, pulmonary, intranasal, buccal, intravenous, intracerebroventricular, intradermal, intramuscular, subcutaneous, intraventricular, intrathecal, intratracheal, intraperitoneal, in utero delivery, or another route of administration or any combination thereof. Other contemplated formulations include projected nanoparticles, liposomal preparations, resealed erythrocytes containing the active ingredient, and i m m u n ogeni c -b ase d formul ati on s .[ 0180| A pharmaceutical composition herein may be prepared, packaged, or sold in bulk, as a single unit dose, or as a plurality of single unit doses. As used herein, a “unit dose” is discrete amount of the pharmaceutical composition comprising a predetermined amount of the active ingredient. The amount of the active ingredient is generally equal to the dosage of the active ingredient which would be administered to a subject or a convenient fraction of such a dosage such as, for example, one-half or one-third of such a dosage.
[0181] The relative amounts of the active ingredient, the pharmaceutically acceptable carrier, and any additional ingredients in a pharmaceutical composition herein may vary, depending upon the identity, size, and condition of the subject treated and further depending upon the route by which the composition is to be administered By way of example, the composition may comprise between 0.1% and 100% (w / w) active ingredient.
[0182] In some embodiments, the pharmaceutical composition comprises a pharmaceutically acceptable carrier and an engineered T-cell receptor herein, a nucleic acid molecule herein and comprising a nucleic acid encoding an engineered T-cell receptor herein, a vector herein comprising the nucleic acid molecule, or a cell herein comprising any of the foregoing
[0183] In addition to the active ingredient, a pharmaceutical composition herein may further comprise one or more additional pharmaceutically active agents.
[0184] As used herein, “parenteral administration” of a pharmaceutical composition includes any route of administration characterized by physical breaching of a tissue of a subject and administration of the pharmaceutical composition through the breach in the tissue ParenteralDOCKET NO. STFD-011-PCT PROVISIONAL PATENT administration thus includes, but is not limited to, administration of a pharmaceutical composition by injection of the composition, by application of the composition through a surgical incision, by application of the composition through a tissue-penetrating non-surgical wound, and the like. In some embodiments, parenteral administration is contemplated to include, but is not limited to, intraocular, intravitreal, subcutaneous, intraperitoneal, in utero delivery, intramuscular, intradermal, intrasternal injection, intratumoral, intravenous, intracerebroventricular and kidney dialyti c infusion techniques
[0185] Formulations of a pharmaceutical composition suitable for parenteral administration comprise the active ingredient combined with a pharmaceutically acceptable carrier, such as sterile water or sterile isotonic saline. Such formulations may be prepared, packaged, or sold in a form suitable for bolus administration or for continuous administration. Injectable formulations may be prepared, packaged, or sold in unit dosage form, such as in ampules or in multi-dose containers containing a preservative. Formulations for parenteral administration include, but are not limited to, suspensions, solutions, emulsions in oily or aqueous vehicles, pastes, and implantable sustained-release or biodegradable formulations. Such formulations may further comprise one or more additional ingredients including, but not limited to, suspending, stabilizing, or dispersing agents. In some embodiments of a formulation for parenteral administration, the active ingredient is provided in dry (i.e. powder or granular) form for reconstitution with a suitable vehicle (e.g. sterile pyrogen-free water) prior to parenteral administration of the reconstituted composition.
[0186] The pharmaceutical compositions may be prepared, packaged, or sold in the form of a sterile injectable aqueous or oily suspension or solution. This suspension or solution may be formulated according to the known art, and may comprise, in addition to the active ingredient, additional ingredients such as the dispersing agents, wetting agents, or suspending agents described herein. Such sterile injectable formulations may be prepared using a non-toxic parenteral ly-acceptable diluent or solvent, such as water or 1,3-butane diol, for example. Other acceptable diluents and solvents include, but are not limited to, Ringer’s solution, isotonic sodium chloride solution, and fixed oils such as synthetic mono- or di-glycerides. Other parentally-DOCKET NO. STFD-011-PCT PROVISIONAL PATENT administrable formulations which are useful include those which comprise the active ingredient in microcrystalline form, in a liposomal preparation, or as a component of a biodegradable polymer systems Compositions for sustained release or implantation may comprise pharmaceutically acceptable polymeric or hydrophobic materials such as an emulsion, an ion exchange resin, a sparingly soluble polymer, or a sparingly soluble salt.[0187} As used herein, “additional ingredients” include, but are not limited to, one or more of the following: excipients; surface active agents; dispersing agents, inert diluents; granulating and disintegrating agents; binding agents; lubricating agents; sweetening agents; flavoring agents; coloring agents; preservatives; physiologically degradable compositions such as gelatin; aqueous vehicles and solvents; oily vehicles and solvents; suspending agents, dispersing or wetting agents, emulsifying agents, demulcents; buffers; salts; thickening agents; fillers; emulsifying agents; antioxidants; antibiotics; antifungal agents, stabilizing agents; and pharmaceutically acceptable polymeric or hydrophobic materials. Other “additional ingredients” which may be included in the pharmaceutical compositions of the invention are known in the art and described, for example in Remington's Pharmaceutical Sciences (1985, Genaro, ed., Mack Publishing Co., Easton, PA), which is incorporated herein by reference.Methods of Treating[0188| In some embodiments, the disclosure relates to a method of treating a hyperproliferative disorder or a viral infection comprising administering a therapeutically effective amount of cells herein to subject in need thereof The cells may comprise a plurality of any cell herein. The cells may be in a pharmaceutical composition herein. In some embodiments, the therapeutically effective amount of cells is about 3xl04cells / kg body weight for patients with less than or equal to 50 kg body weight based on the total cells; e.g., CAR-T cells or NK cells expressing the CAR, in the pharmaceutical composition. In some embodiments, the therapeutically effective amount of cells is about 1.5 x 106cells for patients with more than 50 kg body weight based on the total cells in the pharmaceutical composition. In some embodiments, the therapeutically effective amountDOCKET NO. STFD-011-PCT PROVISIONAL PATENT of cells is from about 3.0 x 10’ to 1.5 x l()bcells. In some embodiments, the therapeutically effective amount of cells is about 10 x 104cells / kg body weight for patients with less than or equal to about 50 kg body weight based on the total cells in the pharmaceutical composition. In some embodiments, the therapeutically effective amount of cells is from about 5 xlO6cells for patients with more than about 50 kg body weight based on the total ceils in the pharmaceutical composition. In some embodiments, the therapeutically effective amount of cells is about 3. Ox 104cells / kg body weight for patients with less than or equal to 50 kg body weight based on the total cells in the pharmaceutical composition. In some embodiments, the therapeutically effective amount of cells is about 1.5 x 10scells for patients with more than 50 kg body weight based on the total cells in the pharmaceutical composition. In some embodiments, the therapeutically effective amount of cells is about IxlO4cells / kg body weight, 2x 104cells / kg body weight, 3 x 104cells / kg body weight, 4 x 104cells / kg body weight, 5 xlO4cells / kg body weight, 6 xl 04cells / kg body weight, 7 xlO4cells / kg body weight, 8 xlO4cells / kg body weight, 9 xlO4cells / kg body weight, 10 xlO4cells / ke bodv weight, 11 x 104cells / kg bodv weight, 12 x 104cells / ke bodv weight, 13 x 104cells / kg body weight, 14 x 104cells / kg body weight, 15 x 104cells / kg body weight, 16 x 1()4cells / kg body weight, 17 x 104cells / kg body weight, 18 x 104cells / kg body weight, 19 x 104cells / kg body weight, 20 x 104cells / kg body weight, 21 x 104cells / kg body weight, 22 x 104cells / kg body weight, 23 x 104cells / kg body weight, 24 x 104cells / kg body weight, 25 x 104cell s / kg body weight, 26 x 104cells / kg body weight, 27 x 104cells / kg body weight, 28 x 104cells / kg body weight, 29 x 104cells / kg body weight, 30 x 104cells / kg body weight, 40 x I04cells / kg body weight, 50 x 104cells / kg body weight, 60 x 104cells / kg body weight, 70 x 104cells / kg body weight, 80 x 104cells / kg body weight, or about 90 x 104cells / kg body weight. In some embodiments, the therapeutically effective amount of cells is about 0.5 x 106cells, 1 xlO6ceils, 1. 5 x 106cells, 2 x 106cells, 2.5 xlO6cells, 3 xlObcells, .3.5 x 106cells, 4 xlO6cells, 4.5 x 10bcells, 5 x 106cells, 5.5 x 106cells, 6 x 106cells, 6.5 x 10° cells, 7 x IO6cells, 7.5 x 106cells, 8 x 10bcells, 8.5 x 106cells, 9 x 106cells, 9.5 x 106cells, 10 x 106cells, 10.5 x 106cells, l l x l 06cells, 11.5 x 106DOCKET NO. STFD-011-PCT PROVISIONAL PATENT cells, 12 x 106cells, 12.5 xlO6cells, 13 x 106cells, 13.5 xlO6cells, 14 xlO6cells, 14.5 xl 00cells, 15 x 10° cells, 16 x 10° cells, 17 x 10° cells, 18 x 10° cells, 19 x 10° cells, 20 x 10° cells, 21 x 10° cells, 22 x 106cells, 23 x J O6cells, 24 x 10° cells, 25 x 106cells, 26 xl 06cells, 27 x 10” cells, 28xl06cells, 29 x 10° cells, 30 x IO6cells, 31 x 106cells, 32 x IO6cells, 33 x 106cells, 34xl06cells, 35 x 106cells, 36 x 106cells, 37 x 106cells, 38 x 106cells, 39 x 106cells, 40 x 10° cells, 41 x 106cells, 42x 10° cells, 43 x 106cells, 44 x 10° cells, or 45 x 106cells.Engineered T-Cells and CARS[01 §9] In some embodiments, the disclosure also relates to an engineered T-cell receptor comprising any one or more antigen or antigenic determinant herein. In some embodiments, the disclosure relates to nucleic acid molecule comprising a nucleic acid sequence encoding an engineered T-cell receptor comprising any one or more antigen or antigenic determinant herein. In some embodiments, the T-cell receptor comprises a plurality of antigens or antigenic determinants herein. Some embodiments comprise a vector comprising the nucleic acid molecule. Further embodiments comprise a cell comprising the engineered T-cell receptor, the nucleic acid molecule, or the vector. Still further embodiments comprise a pharmaceutical composition comprising a pharmaceutically acceptable carrier and the engineered T-cell receptor, the nucleic acid molecule, the vector, or the cell. Yet further embodiments comprise a vaccine comprising the engineered T-cell receptor, the nucleic acid molecule, the vector, the cell, or the pharmaceu ti cal composition ,
[0190] Methods of engineering a T-cell receptor can be found in PCT / US2025 / 035781, which is title “Chimeric Antigen Receptors and Uses Thereof,” was filed June 27, 2025, and is incorporated herein by reference in its entirety. In some embodiments, the vector comprising the nucleic acid encoding a desired engineered T-cell receptor, or CAR, is an adenoviral vector (A5 / 35). In some embodiments, the expression of the nucleic acid sequence(s) encoding a CAR can be accomplished using of transposons such as sleepingDOCKET NO. STFD-011-PCT PROVISIONAL PATENT beauty, crisper, CAS9, and zinc finger nucleases. See below June et al. 2009Nature Reviews Immunology 9. 1 0: 704-716, which is incorporated herein by reference. In brief summary, the expression of a nucleic acid sequence herein encoding an engineered T-cell receptor is typically achieved by operably linking the nucleic acid sequence or portions thereof to a promoter, and incorporating the nucleic acid molecule harboring these elements, a construct, into an expression vector. The vectors can be suitable for replication and integration eukaryotes. Typical cloning vectors that may be implemented in embodiments herein contain transcription and translation terminators, initiation sequences, and promoters useful for regulation of the expression of the desired nucleic acid sequence. The expression constructs of the present disclosure may also be used for nucleic acid immunization and gene therapy, using standard gene delivery' protocols. Methods for gene delivery are known in the art. See, e.g., U.S. Pat Nos 5,399,346, 5,580,859, 5,589,466, which are incorporated by reference herein in their entireties. In some embodiments, the disclosure provides a gene therapy vector.
[0191] The nucleic acid molecule can be cloned into a number of types of vectors. For example, the nucleic acid can be cloned into a vector including, but not limited to a plasmid, a phagemid, a phage derivative, an animal virus, and a cosmid. Vectors of particular interest include expression vectors, replication vectors, probe generation vectors, and sequencing vectors.
[0192] Further, the expression vector may be provided to a cell in the form of a viral vector. Viral vector technology is well known in the art and is described, for example, in Sambrook et al., 2012, MOLECULAR CLONING: A LABORATORY MANUAL, volumes 1 -4, Cold Spring Harbor Press, NY), which is incorporated herein by reference in its entirety, and in other virology and molecular biology manuals. Viruses, which are useful as vectors include, but are not limited to, retroviruses, adenoviruses, adeno- associated viruses, herpes viruses, and lentiviruses. In general, a suitable vector contains an origin of replication functional in at least one organism, a promoter sequence, convenient restriction endonuclease sites, and one or more selectable markers, (e.g., WO 01 / 96584; WODOCKET NO. STFD-011-PCT PROVISIONAL PATENT01 / 29058; and U.S. Pat. No. 6,326,193, which are incorporated herein by reference in their entireties).
[0193] A number of viral based systems have been developed for gene transfer into mammalian cells. For example, retroviruses provide a convenient platform for gene delivery systems. A selected gene can be inserted into a vector and packaged in retroviral particles using techniques known in the art. The recombinant virus can then be isolated and delivered to cells of the subject either in vivo or ex vivo. A number of retroviral systems are known in the art. In some embodiments, adenovirus vectors are used. A number of adenovirus vectors are known in the art. In some embodiments, lentivirus vectors are used.
[0194] Additional promoter elements, e.g., enhancers, regulate the frequency of transcriptional initiation and may be incorporated in a nucleic acid molecule herein. Typically, these are located in the region 30-110 bp upstream of the start site, although a number of promoters have been shown to contain functional elements downstream of the start site as well. The spacing between promoter elements frequently is flexible, so that promoter function is preserved when elements are inverted or moved relative to one another. In the thymidine kinase (tk) promoter, the spacing between promoter elements can be increased to 50 bp apart before activity begins to decline. Depending on the promoter, it appears that individual elements can function either cooperatively or independently to activate transcription. Exemplary promoters of embodiments herein include the CMV IE gene, EF-la, ubiquitin C, or phosphoglycerokinase ( PGK) promoters.
[0195] An example of a promoter of embodiments herein that is capable of expressing a CAR transgene in a mammalian I cell is the EF-1 alpha (EFla) promoter. The native EFla promoter drives expression of the alpha subunit of the elongation factor- 1 complex, which is responsible for the enzymatic delivery of aminoacyl tRNAs to the ribosome. The EFla promoter has been extensively used in mammalian expression plasmids and has been shown to be effective in driving CAR expression from transgenes cloned into a lentivira1 vector. See, e.g., Milone et al., Mol. Ther. 17(8): 1453- 1464 (2009). In some embodiments, the EFla promoter comprises the sequence as known in the art.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0196] In order to assess the expression of a CAR polypeptide or portions thereof, the expression vector to be introduced into a cell can also contain either a selectable marker gene or a reporter gene or both to facilitate identification and selection of expressing cells from the population of cells sought to be transfected or infected through viral vectors. In other embodiments, the selectable marker may be carried on a separate piece of DNA and used in a co- transfection procedure. Both selectable markers and reporter genes may be flanked with appropriate regulatory sequences to enable expression in the host cells. Useful selectable markers include, for example, antibiotic-resistance genes, such as neo and the like
[0197] Methods of introducing and expressing genes into a cell are known in the art and may be implemented in embodiments herein. In the context of an expression vector, the vector can be readily introduced into a host cell, e.g., mammalian, bacterial, yeast, or insect cell by any method in the art. For example, the expression vector can be transferred into a host cell by physical, chemical, or biological means.
[0198] Physical methods for introducing a polynucleotide into a host cell include calcium phosphate precipitation, lipofection, particle bombardment, microinjection, electroporation, and the like. Methods for producing cells comprising vectors and / or exogenous nucleic acids are well-known in the art. See, for example, Sambrook et al., 2012, MOLECULAR CLONING: A LABORATORY MANUAL, volumes 1 -4, Cold Spring Harbor Press, NY). .Another method for the introduction of a polynucleotide into a host cell is calcium phosphate transfection.
[0199] Biological methods for introducing a polynucleotide of interest into a host cell include the use of DNA and RNA vectors. Viral vectors, and especially retroviral vectors, have become the most widely used method for inserting genes into mammalian, e.g., human cells. Other viral vectors can be derived from lentivirus, poxviruses, herpes simplex virus 1 , adenoviruses and adeno-associated viruses, and the like. See, for example, U.S. Pat. Nos. 5,350,674 and 5,585,362, which are incorporated herein by reference in their entireties.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0200] Regardless of the method used to introduce exogenous nucleic acids into a host cell, in order to confirm the presence of the recombinant DNA sequence in the host cell, a variety of assays may be performed. Such assays include, for example, “molecular biological” assays well known to those of skill in the art, such as Southern and Northern blotting, RT- PCR and PCR, “biochemical” assays, such as detecting the presence or absence of a particular peptide, e.g., by immunological means (ELISAs and Western blots) or by assays described herein to identify agents falling within the scope of the invention.
[0201] The vector comprising an engineered T-cell receptor, or CAR, herein can, in some embodiments, be directly transduced into a cell, e.g., a T cell or NK cell. In some embodiments, the vector is a cloning or expression vector, e.g., a vector including, but not limited to, one or more plasmids (e.g., expression plasmids, cloning vectors, minicircles, minivectors, double minute chromosomes), retroviral and lentiviral vector constructs. In some embodiments, the vector is capable of expressing the CAR construct in mammalian T cells or NK cells. In some embodiments, the mammalian T cell is a human T cell.
[0202] Any and all journal articles, patent applications, issued patents, or other cited references disclosed herein are incorporated by reference in their respective entireties.References1. Pam er, E. & Cresswell, P. Mechanisms of MHC class I— restricted antigen processing. Annu Rev Immunol 16, 323-358 (1998).2. Neefjes, J., Jongsma, M. I... M,, Paul, P. & Bakke, O. Towards a systems understanding of MHC class I and MHC class II antigen presentation. Nat Rev Immunol 11, 823-836 (2011).3. Shiina, T., Hosomichi, K., Inoko, H. & Kulski, J. K. The HLA genomic loci map: expression, interaction, diversity and disease. J Hum Genet 54, 15 -39 (2009).4. Bassani-Stemberg, M. & Gfeller, D. Unsupervised HLA Peptidome Deconvolution Improves Ligand Prediction Accuracy and Predicts Cooperative Effects in Peptide-HLA Interactions. J Immunol 197, 2492-2499 (2016).DOCKET NO. STFD-011-PCT PROVISIONAL PATENT5. Pearson, H. et al. MHC class Lassociated peptides derive from selective regions of the human genome. J Clin Invest 126, 4690-4701 (2016).6. McCutcheon, J. A., Gumperz, J., Smith, K 822 . D., Lutz, C. T. & Parham, P. Low HLA-C expression at cell surfaces correlates with increased turnover of heavy chain mRNA. J Exp Med 181, 2085-2095 (1995).7. Jurtz, V. et al. NetMHCpan-4.0: Improved Peptide-MHC Class I Interaction Predictions Integrating Eluted Ligand and Peptide Binding Affinity Data J Immunol 199, 3360-3368 (2017)8. Olivier, T., Haslam, A., Tuia, J. & Prasad, V. Eligibility for Human Leukocyte Antigen-Based Therapeutics by Race and Ethnicity. JAMANetw Open 6, e2338612 (2023).9. Schmittel, A , Keilholz, U. & Scheibenbogen, C. Evaluation of the interferon -gamma ELISPOT- assay for quantification of peptide specific T lymphocytes from peripheral blood. J Immunol Methods 210, 167-174 (1997).10. Darwish, M. et al. High-throughput identification of conditional MHCI ligands and scaled up production of conditional MHCI complexes. Protein Sci 30, 1169-1183 (2021).11. Haj, .A. K. et al. High-Throughput Identification of MHC Class I Binding Peptides Using an Ultradense Peptide Array. J Immunol 204, 1689-1696 (2020).12. Gurung, H. R. et al. Systematic discovery of neoepitope-HLA pairs for neoantigens shared among patients and tumor types. Nat Biotechnol (2023) doi:10.1038 / s41587-023-01945-y.13. Bruno, P. M. et. al. High-throughput, targeted MHC class I immunopeptidomics using a functional genetics screening platform. Nat Biotechnol 41, 980-992 (2023).14. Greten, T. F. et al. Peptide-p2~microglobulin-MIIC fusion molecules bind antigen-specific T cells and can be used for multivalent MHC-Ig complexes. Journal of Immunological Methods 271, 125-135 (2002).15, Yu, Y. Y. L., Netuschil, N., Lybarger, L., Connolly, J. M. & Hansen, T. H. Cutting Edge: SingleChain Trimers of MHC Class I Molecules Form Stable Structures That Potently Stimulate Antigen- Specific T Cells and B Cellsl . The Journal of Immunology 168, 3145 847 3149 (2002).16. Truscott, S. M. et al. Disulfide Bond Engineering to Trap Peptides in the MHC Class I Binding Groovel . The Journal of Immunology 178, 6280-6289 (2007)DOCKET NO. STFD-011-PCT PROVISIONAL PATENT17. Blum, J. S., Wearsch, P. A. & Cresswell, P. Pathways of Antigen Processing. Annu Rev Immunol 31, 443—473 (2013).18. Pishesha, N., Harmand, T. J. & Ploegh, H. I... A guide 852 to antigen processing and presentation. Nat Rev Immunol 22, 751-764 (2022).19. Reynisson, B., Alvarez, B., Paul, S., Peters, B. & Nielsen, M. NetMHCpan-4. 1 and NetMHCIIpan-4.0: improved predictions of MHC antigen presentation by concurrent motif deconvolution and integration of MS MHC eluted ligand data. Nucleic Acids Research 48, W449- W454 (2020).20. As, S. et al. SARS-CoV-2 Epitopes Are Recognized by a Public and Diverse Repertoire of Human T Cell Receptors. Immunity 53, (2020).21. Sekine, T. et al. Robust T Cell Immunity in Convalescent Individuals with Asymptomatic or Mild COVID-19. Cell 183, 158-168.e14 (2020)22. Tate, J. G. et al. COSMIC: the Catalogue Of Somatic Mutations In Cancer. Nucleic Acids .Research 47, D941-D947 (2019).23. Sarkizova, S. et al. A large peptidome dataset improves HLA class I epitope prediction across most of the human population. Nat Biotechnol 38, 199-209 (2020).24. Smyth, R. P et al. Reducing chimera formation during PCR amplification to ensure accurate genotyping Gene 469, 45-51 (2010).25. Jetzt, A. E. et al. High Rate of Recombination throughout the Human Immunodeficiency Virus Type 1 Genome. J Virol 74, 1234-1240 (2000).26. Tran, E et al T-Cell Transfer Therapy Targeting Mutant KRAS in Cancer. N Engl J Med
[0188] 871 375, 2255-2262 (2016).27. Choi, J. et al. Systematic discovery and validation of T cell targets directed against oncogenic KRAS mutations. Cell Rep Methods 1 , 100084 (2021).28. Gonzalez-Galarza, F. F. et al. Allele frequency net database (AFND) 2020 update: gold standard data classification, open access genotype data and new query tools. Nucleic Acids Res 48, D783-D788 (2020).DOCKET NO. STFD-011-PCT PROVISIONAL PATENT29. Gfeller, D. & Bassani -Sternberg, M. Predicting Antigen Presentation-— What Could We Learn From a Million Peptides? Front Immunol 9, 1716 (2018).30. Karnaukhov, V. et al HLA variants have different preferences to present proteins with specific molecular functions which are complemented in frequent haplotypes. Frontiers in Immunology 13, (2022).31. Sidney, J., Peters, B., Frahm, N., Brander, C. & Sette, A. H 882 LA class I supertypes: a revised and updated classification. BMC Immunology 9, 1 (2008),32. Sarkizova, S. et al. A large peptidome dataset improves HLA class I epitope prediction across most of the human population. Nat Biotechnol 38, 199-209 (2.020).33. Latysheva, N. S. & Babu, M. M Discovering and understanding oncogenic gene fusions through data intensive computational approaches. Nucleic Acids Res 44, 4487-4503 (2016).34. Liu, Y, et al Etiology of oncogenic fusions in 5, 190 childhood cancers and its clinical and therapeutic implication. Nat Commun 14, 1739 (2023).35 Zhang, H. et al. Identification of NY-ESO-1157-165 Specific Murine T Cell Receptors With Distinct Recognition Pattern for Tumor Immunotherapy. Front Immunol 12, 644520 (2021).36. Wooldridge, L. et al. A Single Autoimmune T Cell Receptor Recognizes More Than a Million Different Peptides. J Biol Chem 287, 1168-1177 (2012)37. Birnbaum, M, E. et al. Deconstructing the peptide-MHC specificity of T cell recognition. Cell 157, 1073-1087 (2014).38. Giannakopoulou, E. et al. A T cell receptor targeting a recurrent driver mutation in FLT3 mediates elimination of primary human acute myeloid leukemia in vivo. Nat Cancer 4, 1474-1490 (2023).39. Yamada, T. et al. EGFR T790M mutation as a possible target for immunotherapy; identification of HLA-A* 0201 -restricted T cell epitopes derived from the EGFR T790M mutation. PLoS One 8, e78389 (2013).40. Robinson, J. et al. The IPD and IMGT / HLA database: allele variant databases. Nucleic Acids Res 43, D423-431 (2015).DOCKET NO. STFD-011-PCT PROVISIONAL PATENT41. Hickman, H. D. et al. Toward a definition of self proteomic evaluation of the class I peptide repertoire. J Immunol 172, 2944-2952 (2004).42. Sachs, A. et al Impact of Cysteine Residues on MHC Binding Predictions and Recognition by Tumor-Reactive T Cells. J Immunol 205, 539-549 (2020).43. Kennedy, P. R., Barthen, C., Williamson, D. J. & Davis, D. M. HLA-B and HLA-C Differ in Their Nanoscale Organization at Cell Surfaces. Front Immunol 10, 61 (2019).44. Pearlman, A. H. et al. Targeting public neoantigens for cancer immunotherapy Nat Cancer 2, 487-497 (2021).45. Yu, B. et al. Engineered cell entry links receptor biology with single-912 cell genomics. Cell 185, 4904-4920. e22 (2022).46. Kula, T. et al. T-Scan: A Genome-wide Method for the Systematic Discovery of T Cell Epitopes. Cell 178, 1016-1028. e!3 (2019).EXAMPLESExample 1: Massively parallel immunopeptidome by DNA sequencing reveals landscape of antigen presentation
[0189] An engineered single-chain construct that allows modular and combinatorial construction of peptide x HLA alleles for massively parallel testing of binding affinity. A cell based assay reads out greater than75000 combinations at once by DNA sequencing. This method is called ESCAPE- seq. We generated the most comprehensive map of cancer antigen presentation that is applicable to 90% of the world’s population. A list of cancer antigen epitopes that most broadly presented across different HLA alleles has been identified. See Table C and C.1.
[0190] ESCAPE-seq (Enhanced Single-chain Antigen Presentation sequencing) is also described herein. ESCAPE-seq is a massively parallel platform for comprehensive screening of Class I HLA- peptide combinations for antigen presentation via deep DNA sequencing. ESCAPE-seq demonstrates flexibility, high throughput, sensitivity, and ease of implementation. Its capability is showcased in simultaneously assessing > 75,000 peptide-HLA combinations, including understudied HLA-C alleles, identifying immunogenic regions in SARS-CoV-2 variants, andDOCKET NO. STFD-011-PCT PROVISIONAL PATENT broadly presented oncogenic driver mutations and fusions across diverse HLA alleles. ESCAPE- seq is a promising tool for high-throughput antigen presentation discovery, offering insights into HLA. diversity and potential avenues for improving immunotherapy
[0191] Among other applications, ESCAPE-seq may be utilized for epitope discovery' for vaccines against cancer or infectious diseases, epitope identification for immunomodulation, TCR-T cell design, or T cell engineering.Methods
[0192] DNA synthesis and plasmid construction.
[0193] All plasmids were made with Gibson assembly (NEB) unless specified otherwise. Briefly, based on lentiviral vector of single-chain trimer pHLA-A0201 and pHLA-A010145, P2A- puromysin-2A-eGFP was inserted after cytoplasmic tail of HLA alleles. Then oligos encoding different peptides were inserted into SCT between a signal peptide and B2m gene with a flexible linker. For building SCT trimer of other MHC alleles, we made the Y84C mutation first, then insert to replace A *02 or A*01 above. For example, the HLA- A*02 was digested out and replaced with B*0702 (addgene, #135509), C*0401 and H2kb allele (synthesized by Twistbio).
[0194] Oligo encoding Sequences encoding human HLA Alleles were obtained from IEDB. HLA- A0101, HLA-A0201, and HLA-B0702 For each allele, a cloning lentiviral vector was made with Esp3i sites in the place of peptide in the single-chain format45, followed by a P2A sequence, eGFP, T2A and puromycin resistance gene. Afterward, various peptides obtained from literature or IEDB database (Tables 1 and 2) were inserted in place.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0195] For HLA-eGFP direct fusion, the P2A sequence between HLA and eGFP above were placed with a flexible GS rich linker. Various point mutants and deletion mutants were generated similarly using the above vector as source. To express peptides alone in the endoplasmic reticulum (ER), a CM V promoter and the signal peptide from the human growth hormone gene were utilized. The peptide expression was driven in the ER using a lentiviral vector that included blasticidin resistance. All pooled oligo encoding peptides were ordered from TwistBio. For benchmark against IEDB database, IEDB database was retrieved, and all peptides for 4 common HLA alleles were extracted, A*01()l, A*0201, B*0702, and C*0401. Then peptides were randomly selected across all affinities for each allele and combined to make a pool of > 2000 peptide in total. Then the oligo pool encoding the peptides flanked by overhang sequences on both sides was inserted into theDOCKET NO. STFD-011-PCT PROVISIONAL PATENT cloning vector as described above and was electroporated into competent cells (Enduro electrocompetent cells, Biosearch Technologies) using Biorad MicroPulser (Biorad). For screening on SARS-CoV2, DNA sequences encoding Spike and protein N from strain Wuhan-Hu- 1 were used. A tiling pool was made to cover full sequencing with 3bp shift in frame between neighbor tiles. All strain variants peptides were included in the pool based on strain variants data (The GFF file of Dec 2021 from uniprot).
[0196] C ombinatorial pools generation
[0197] 50 HLA Class I alleles were selected based on their high frequencies in population in diverse world population and races (https: / / www.allelefrequencies.net / hla.asp) or associated with pathogenic diseases. Overall 14 HLA- A, 29 HLA-B and 7 HLA-C were obtained accordingly, which are A*01 :01, A*02:01, A*02:05, A*02:12, A*03:01, A*ll:01, A*23:01, , A*24:02, .4*30:01, A*31 :01, A*31:08, A*34:01, A*33:03, A*68:01, B*07:02, B*08:01, B*08:02, B*13:01, B* 15:01, B*15:02, B*15:03, B*18:01, 8*27:01, B*27:05, B*27:02, B*35:01, B*35:02, B*35:08, B*39:06, B*40:01, B* 40:06, B*40:10, B*41 :01, B*44:0, B*46:01, B*48:03, B*50:01 , B*5l :0l, B*52:01, 8*54:01, B*56:01, B*57:01, B*58:01, ('*03:04, C*04:01, ('*06:02. C*07:01, ('*07:02. C*08:02, C* 12:03. Sequences encoding human HLA Alleles were obtained from IEDB. Some restriction enzyme sites if any in HLA were mutated without changing amino acid and Y84C mutation introduced. Then part of b2m sequence, a unique barcode per HLA alleles and a long linker containing the 10x TSO sequence were added to HLA sequence and synthesized by Twist Bio. An oligo pool of over 1500 peptides encoding cancer driver gene peptides and oncogenic fusion gene peptides was collected as described below. Then flanking sequences for Gibson assembly were added to either side of each oligo and synthesized by Twist Bio. The combinatorial paired peptide-HLA plasmid pool can be generated by 2 steps. First, the 50 HLA alleles gene fragments were amplified by PCR with very low cycler number to get to yield about ~ 50ng as we noticed that PCR with high cycle number produces chimeric HLA alleles due to its high homology between each other. Alternatively, HLA alleles fragments can be inserted to lentiviral vector via digestion and ligation without PCR step; then inserted with peptides pool with Gibson reaction. This avoids the PCR step that potentially generates chimeric fragments. In this study, each HI..A-1DOCKET NO. STFD-011-PCT PROVISIONAL PATENT allele was inserted individually to generate 50 cloning vectors. Then the plasmids were pooled to insert peptide pools as described below. This strategy offers flexibility of customized selection ofHLA alleles for different peptides pool of interest in the future[01981 After that, peptide oligo pools were amplified 6 cycles with PCR. The cloning plasmid pool with 50 HLA inserted (as described above) was digested with Esp3i and the peptide pool was inserted to it via Gibson to generate the final combinatorial pHLA pooled plasmid. The aim was to get over 50k colonies for first Gibson of 50 HLAs, while the last pl ILA pool should give colonies of >100x (lOOOx ideally, between 7.5-75M colonies) of library’s complexity Also generated were spike-in pools to mix during the experiments. To examine the chimeric and recombination of pHLA reads, two spike-in pools were generated separately, with each consisting of 25 HLA alleles with their corresponding peptide pools (Table 3).DOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0199] Cell lines generation and culture
[0200] HEK293T cells were cultured in DMEM supplemented with 10% FBS and 1% Penstrep. HL A knockout using cas9 RNP was done as described previously. sgRNA were synthesized by Synthego including ACUGCUACUUCUCGCCGACU (Human TAPI) (SEQ ID NO: 1), CUGGUGGGGUACGGGGCUGC (human TAP2) (SEQ ID NO: 2), CGGCUACUACAACCAGAGCG (HLAJ) (SEQ ID NO: 3),AGAUCACACUGACCUGGCAG (HL. A-2) (SEQ ID NO: 4), AGGUCAGUGUGAUCUCCGCA (HLA-3) (SEQ ID NO: 5). For dual HLAand TAP knock-out cells, HLA-KO HEK203T cells were used to electroporate cas9 RNP mixed with TAP1 / TAP2 sgRNA. The cells were cultured for 5 days. Then the cells were stained with PE-b2m (Biolegend) and were sorted into single cells seeded into 96-well plate. The clonal cell wells were picked and expanded. For each well, its TAP 1 / 2 were verified with genomic PCR and sanger sequence,
[0201] Lentivirus production and titration
[0202] Lentivirus were made as described previously. Briefly, per 6-well, HEK293T were transfected with a viral expression vector (2ug), pMD2.G (VSV-G wiki type) (lug), and psPax2 (2ug) with lipofectamine 3000. The media was changed one next day, and viral supernatants were collected twice at 48 hr and 72hr respectively. The virus was concentrated with 4x Lenti-X according to manufacturer’s protocol and stored at 20x concentrated in -80DC. For pooled pHLA vims, either 6cm or 10cm dishes were used to make a larger quantity of virus, with proportional scaling in DNA and reagent amount. The viruses were first titrated with HEK293T cells at 25% confluence, and percentage of infection is measured by a flow cytometry (Attune).
[0203] Transfection, cell assay and flow cytometry'
[0204] 100K cells were seeded onto 24-well plates and cultured overnight. Next day, 0.75ul DNA with 1.5ul lipofectamine 3000 (Thermofisher) was transfected into cells. The cells were collected 1 or 2 days after transfection, incubated in full DMEM media for 30min at 37c to recover. Afterward, the cells were pelleted and stained with 2ul of antibodies (PE anti-b2m from Biolegend,DOCKET NO. STFD-011-PCT PROVISIONAL PATENTPe-Cy7 anti human HLA-A2 from Biolegend ) for 30min on ice. Cells were washed once before examined with a flow cytometry (Attune, Lifetechnology).
[0205] Imaging
[0206] Cells were transfected with HLA-eGFP fusion constructed as described above. The images were taken with Zeiss LSM780 confocal microscope next day.
[0207] Pooled antigen presentation screening
[0208] The cells were infected with virus at MOI at - 0.15. Puromycin was added to cells after 2 days of infection at 2ug / ml final. The cells were collected after 4-6 days of drug selection. After incubating in fresh DMEM media for 30min at 37c, the cells were stained with 2ul of anti-b2m- PE (Biolegend) at 2ul per 1 million of cells on ice for 30min with intermittently mixing. Once washed, the cells were resuspended and sorted into 4 fractions based on PE-b2m intensity by BD Aria. We typically require 2 biological replicates per screen and starting cells about > 1500x times of the peptide pool’s complexity for each replicate.
[0209] DNA extraction, library generation and sequencing
[0210] Sorted cells were spun down and genomic DNA were extracted with Zymo column (Zymo QuickDNA) according to manufacturer’s protocol, or with lysis and precipitation. Briefly, the cells were first resuspended in lysis buffer (20m M Tris, 5mM EDTA and 50mM NaCl., 0.1% SDS), then Rnase A and proteinase K (20mg / ml stock) were added at 5ul per lOOul solution. The samples were incubated at 37c for 3()min then 50c overnight Next day, aqueous phase containing the DNA was obtained using Phenol:Chloroform:Isoamyl Alcohol (Invitrogen) and Maxtract High Density from Qiagen following manufacture’s protocol. DNA was precipitated with isopropanol 70% following standard protocol.
[0211] The library that encodes HLA peptide and barcodes were generated through 3 round of PCR (with primers in Table 4).DOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0212] First, the pHLA fragment was enriched from genomic DNAby 15 cycles of PCR (98°C for 3 min, then 15x of 20 sec at 98°C 20 sec, 20 sec at 58°C and 60sec at 72°C) with 0.8uM primers of amp GH Fw, and b2m-bc rev. After cleanup, 5ul of elution was used the second round of PCR with 0.8um nested primer containing Illumina adapters P7 GH HLA fw and P5 BC- B2m rev (98°C for Imin, then 6x of 20 sec at 98°C 20 sec, 20 sec at 59°C and 60sec at 72°C). The aboveDOCKET NO. STFD-011-PCT PROVISIONAL PATENT primers were designed in a way compatible with dual index used in Illumina sequencing platforms The final libraries were obtained by 3rd round of index PCR with Illumina Truseq based index primers (98°C for Imin, then 6x of 20 sec at 98°C 20 sec, 20 sec at 63°C and 60sec at 72°C). Due to low complexity of the libraries, typically 25% phix were spiked in when sequencing by Nextseq or Hiseq.
[0213] Database retrieve and collection
[0214] Three distinct sets of peptides were created for the evaluation. For IEDB-IC50, 2178 peptides with a broad range of measured IC50s were selected from the IEDB for alleles HLA- A0*l :01, HLA-A*02:01, HLA-B*07:02, HLA-C*04:01. For HLA-C, due to the scarcity of strong-binding measurements, a smaller selection of 150 peptides was opted for due to the lack of existing measured IC50s. The E-scores of these peptides were measured in four separate experiments, one for each allele
[0215] T he sequence of SARS-CoV2 Wuhan- 1 strain was used as wild-type. The GFF file containing different strain mutation was obtained from uniprot (accession: P0DTC2, as of 2021), from which the list of mutations per strain were extracted to build the peptide pools. In this study for simplicity, only the point mutations and in- frame deletion / mutation in Spike protein and protein N of SARS-CoV2 (Table 5) were considered, which give > 26009-mers antigens for alleles HLA-A*01:01, HLA-A*02:01, HLA-B *07:02, Their E-scores were measured concurrently with the IEDB experiments.
[0216] Top cancer driver gene mutations were collected from COSMIC database. Briefly, oncogenes in the database were first manually ranked based on occurrence. Then the top prevalent point mutations per genes were collected to get a candidate list. DNA fragments encoding peptides across each mutation were extracted from gene sequences from ncbi website
[0217] For fusion mutations, cosmic fusion genes were similarly ranked based on mutation numbers from COSMIC database and picked top - 30 fusion variants (e.g. >5-10% or occurrence >50). Per variant, the genomic coordinates for the break point (BP) were retrieved first. Then, fusion genes” DNA and Protein sequences were obtained based on genomic coordinates using FusionGDB2 (https: / / corapbio.uth.edu / FusionGDB2 / index.html). Next, ucsc genome browser wasDOCKET NO. STFD-011-PCT PROVISIONAL PATENT used to find the DNA sequence for these two genes (version hgl 9), align to the DNA sequence to find BP site on the coding sequence. The BP per fusion was further confirmed by Uniprot (htps: / / www.uniprot.org / uniprotkb), from where protein sequences for the two genes were aligned to fusion protein sequence obtained above to retrieve peptides across BP. Per mutation / fusion point, a pool of the peptides were obtained by tiling 9—11 amino acid short peptides across the mutation or breakage point. Altogether we tested 1500 top oncogenes from the COSMIC database across 50 diverse alleles from HLA-A, HLA-B and HLA-C.
[0218] Data analysis
[0219] Custom python scripts were built to read the fastq files using Biopython functions and count the paired reads. Here we only took the antigen peptides or HLA barcodes with exact match of their flanking sequence between 6-10bp. Then a count table containing read counts for each peptide-HLA pairs per sorted bin was built.
[0220] E-score calculation and clustering
[0221] The transformation of four binned numbers into an E-score commences with a normalization of all the reads within each bin through dividing the counts in each bin bv the average. This standard procedure in genomic studies was performed to account for any discrepancies that may arise due to variations in sequencing read depth Differentiating from the conventional RNA-seq read normalization, we omitted log normalization in our process. Our rationale stems from the presumption that the distributions of our counts are unlikely to conform to a log-normal distribution, thus log normalization may not provide an accurate representation of our data. Subsequently, we normalized the total number of sequencing reads measured per trimer (pHLApair). This step is important in adjusting for uneven distribution of antigen-HLA reads that were originated from various steps in the ESCAPE-seq experiments, including pooled plasmid cloning, virus production and transduction, and library construction. With the normalization complete, each allele-peptide-trimer now possesses normalized counts, indicating the number of cells detected in each respective bin. To synthesize these data into a singular E-score, we used the formula: Escore = counts i,g* Whg + countsi™ * wi0W+ countSmed * Wmed + countSh^ * Whtgh, whereDOCKET NO. STFD-011-PCT PROVISIONAL PATENTM’fg = 0, w / ov,- 2, w!!!ed:::4, Whlsh:r:5, where we put weight at log scale that matched the binning scale during cell sorting, while assigning background bin as 0.
[0222] For combinatorial screening with multiple HLA. alleles, we further normalized E- score for each allele to align their distributions in a more comparable manner. This normalization per allele is instrumental for the alignment of our data io a specific reference. This is to mitigate the influence of varying allele efficacies in presentation and / or its intrinsic stability to escape from ER. To facilitate this, we calculated the mode of lower peak (FIG 9F) for each allele (presumably from all negative peptides). Following this, we adjusted every E-score for that, specific allele such that these modes converge This technique of alignment ensured that the E-score is not skewed due to the intrinsic variations between different alleles, thus making the comparison between different alleles more accurate and meaningful.
[0223] Heatmaps were generated using the Corapl exHeatmap R package. Only peptides with at least 1 HLA E-score above threshold were included. For the heatmap in Fig 4, the score matrix was clipped to values between 2 and 5 to enhance visualization.
[0224] Evaluation Metrics
[0225] NetMHC predictions for individual peptides in binding affinity mode ( ba) and elutation mode (__el) were obtained through NetMHC 4.1”s web interface (peptide mode in https: / / services.healthtech.dtu.dk / services / NetMHCpan-4, 1 / ). Predictions for pooled peptides were obtained through command line version of NetMHC 4.1 (downloaded from https: / / downloads.iedb.org / tool s / ) .
[0226] When contrasting the measured IC50s, E-scores, and NetMHCpan predictions, we employed both regression metrics (Pearson’s r and Spearman’s r) and classification metrics (ROC- AUC and PR-AUC), calculated separately for each allele. The regression analyses were performed after transforming IC50s into a log* transformed version, following the format used in training predictive models: 1 - log(binding affinity ) / log(50, 000) as detailed in source. This transformation resulted in affinity values in the range from 0 to 1, with an IC50 greater than 0.426 corresponding to an IC50 less than 500nM. Concurrently, for classification analyses involving ROC-AUC and PR-AUC, a thresholding of the labels was implemented For IC50-based labels, including thoseDOCKET NO. STFD-011-PCT PROVISIONAL PATENT from IEDB measurements and NetMHCpan predictions, binders were defined as those with IC50 values below 500nM. Specifically, Pearson and Spearman correlation were calculated using module spearmanr and pearsonr from scipy .stats. ROC and PRC and their AUC were calculated functions in python package skleam. metrics. To get the error bars, we visualized the 95% confidence interval of the metrics across all the alleles. For labels derived from ESCAPE measurements in this paper, binders were classified as entities with an E-score greater than 3.2 for single HL A allele screening. For combinatorial ESCAPE-seq, we raised the cutoff to 3.8 due to higher observed noisy, presumably from multiple HLA alleles. In general, a cut-off between 3.5-4 is reasonable.
[0227] Furthermore, for each metric calculated, we established a 95% confidence interval. This was achieved through the utilization of a non-parametric bootstrap method, entailing the generation of 1000 samples from the original dataset with replacement, thereby offering an estimation of the uncertainty inherent in our metric estimates.
[0228] Evaluating predictions of E-scores by NetMHCpan models
[0229] To ascertain the degree of alignment between current state-of-the-art model predictions and E-scores, we calculated metrics using NetMHCpan-4.1 predictions as an estimation for E-scores, employing a threshold of E-score > 3.8 for the classification metrics. NetMHCpan-4.1 includes two model variants: a binding-affinity (BA) model trained to predict IC50s and an eluted-ligand (EL) model trained to predict ligands identified on the cell surface via Mass Spectrometry Although the EL model is generally the standard, we observed a superior correlation with E-scores in the BA model, which we therefore adopted for comparative metrics.
[0230] All metrics indicated a high degree of congruence between model predictions and E-scores when utilizing the IEDB data, which formed the training basis for these models. However, when the models were applied to the SARS and oncogene datasets, the metrics were less encouraging. Besides ROC-AUC, which is largely impervious to class imbalances characteristic of the datasets we analyzed, there was a pronounced discrepancy in performance between the IEDB metrics and those derived from the SARS and oncogene datasets.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0231] Previous studies have demonstrated a correlation between the number of peptides in the training data for a given allele and the performance of the model on other peptides of the same allele However, in our analysis, the number of training examples for a given allele failed to account for most of the variance observed in the model’s inaccurate prediction of stability in the oncogene and SARS datasets This was despite the model demonstrating competent performance when evaluated on the same alleles using peptides from the training set. While a slight dip in performance is predictable when evaluating peptides absent from the training data, a 30-75% performance reduction across most metrics implies an increased challenge posed by these new datasets. The smaller proportion of positive peptides per allele in these datasets could contribute to this difficulty. However, even when resampling the data to match the percentage of positives in the IEDB dataset, a significant gap in prediction ability persisted. This observed drop-off in performance underpins the necessity of the ESCAPE methodology to augment current computational prediction methods and address the identified performance gaps.
[0232] Exploring allele similarity as defined by peptide-preferences
[0233] To evaluate the similarity among HLA alleles, we may consider three different parameters: sequence similarity, structural similarity, or functional similarity (the similarity in their binder peptides). Although crystal structure information is not available for all HLA alleles, sequence data are available, and with the introduction of our new oncogene dataset, we can compare the binding preferences of 50 distinct alleles. The sequences used to compare alleles are constituted of 34 amino acids from each allele, often denoted as MHC pseudo-sequences. These sequences encapsulate all the amino acids within 5 A of the peptide binding cleft that vary between alleles We calculate sequence similarity using the equation:where AASim(ai, a.2) represents the similarity between amino acids ai and a.2 as per the BLOSUM62 matrix.
[0234] To quantify the similarity among alleles based on peptide preferences, we consider the list of E-scores for all 1500 peptides in the oncogene pool and calculate the cosine distance betweenDOCKET NO. STFD-011-PCT PROVISIONAL PATENT the E-scores for all peptides between each allele pair. These cosine di stances are used for clustering the alleles and for visualizing alleles in two dimensions. Specifically, UMAP was constructs with python umap package(with parameters n components:::2, random shite 42. n neighbors :v min__dist=0.0L metric- 'cosine" ).
[0235] T Cell Immunogenicity and T cell reporter Assays
[0236] .Antigen-specific T cell immunogenicity was evaluated by a previously publishedprotocol46. Briefly, HLA-typed PBMCs from 4 healthy donors were utilized (C.T.L.).Cryopreserved PBMCs were quickly thawed in 37°C water bath and transferred into RPMI medium (Thermo Fisher 815 Scientific) containing DNase I (Sigma-Aldrich) at a final concentration of 2 U / mL, spun down and resuspended in X-VIVO 15 medium (Lonza) supplemented with cytokines promoting dendritic cell (DC) differentiation, GM939CSF (Peprotech, 1000 lU / mL), IL-4 (R&D Systems, 500 lU / mL) and Flt3L (Peprotech, 50 ng / mL). Cells were seeded at 105 cells per well in U bottom 96-well plates and cultured for 24 hours before being stimulated with control reagents 941 or pooled test peptides (custom peptide synthesis, JPT Peptide Technologies), where each peptide was at a final concentration of 1 uM, together with adjuvants promoting DC maturation, LPS (Invivogen, 0.1 mg / mL), R848 (Invivogen, 10 mM) and IL- 1 P (R&D Systems lOng / mL), in X-VIVO 15 medium. Starting 24 hours after stimulation, cells were fed every2-3 days with cytokines supporting T cell expansion, IL-2 (R&D Systems, 10 IIJ / mL), IL-7 (Peprotech, 10 ng / mL) and IL-15 (Peprotech, 10 ng / mL) in complete RPMI media (GIBCO) containing 10% human serum (R10). After 10 days of culture, cells were harvested, pooled within groups, washed, resuspended in R 10 and seeded at 2x105cells / well in U-bottom 96- well plates. Expanded T cells were then re-stimulated with control reagents or 1 p.M of test peptides, either pooled or individual, together with 0.5mg / mL of costimulatory antibodies, anti- CD28 (BD Biosciences) and anti-CD49d (BDBiosciences), and protein transport inhibitors BD Golgi Stop TM, containing monensin and BD GolgiPlugTM, containing brefeldin A, utilized at manufacturer’s recommended concentrations. After 8 hours of incubation at 37°C, cells were processed for intracellular staining for flow cytometry using BD Cytofix / CytopermTM reagents according to manufacturer’s protocol. The following antibodies were used: for surface stainingDOCKET NO. STFD-011-PCT PROVISIONAL PATENTCD3 (clone SK7, F1TC), CD4 (done RPA-T4, BV785) and CD8a (done RPA959T8, APC) and for intracellular staining IFN-g (done B27, PE) and TNF-a (done Mab 11,PE / Cy7). All antibodies were purchased from BioLegend. LIVE / DEAD Fixable AquaDead Cell Stain Kit (Thermo Fischer Scientific) was used for live and dead cell discrimination. Data was acquired using the Invitrogen Attune NxT flow cytometer and FlowJo V10 was used for analysis. DMSO (Sigma-Aldrich) was used at the equal volume of the test peptides and served as the vehicle / negative control. Significance was evaluated by t test comparing DMSO vs peptide-specific cytokine formation by CD8+ Tcells. For T cell reporter assay in Figure S 1 C, we transfected HLA / TAP KO HEK293Tcells with SCTs containing either HLA-A2:SLLMWITQC (SEQ ID NO: 1333) (NY-ESO-1 epitope) or HLA968A2:NLVPMVATV (SEQ ID NO: 1304) (CMV epitope). After 48h, SCT transfected cells were cocultured with Jurkat NFAT-GFP reporter cells stably expressing NY-ESO-1 -specific 1G4 TCR. TCRactivation was measured by quantifying GFP+ Jurkat cells using flow cytometry (BDFACSCanto) at coculture time points ranging from 5 min to 24 hours.
[0237] Mass Spectrometry
[0238] MS experiments based on HLA-C*0304 allele were done according to previouslypublished16. We first transduced wild-type HLA-C*0304 allele with lentivirus with puromycin resistance into HEK293T cells with dual HLA and TAP knock-out as described above. Afterward, a pool of peptides with a range of low and high E-score with HLA-C*0304 from ESCAPE-seq were transduced to the cells with lentivirus carrying blasticidin resistance. Cells were selected with puromycin and blasticidin and expanded for 2 weeks. - 600M of cells were collected for MS experiments. MS experiments were outsourced to MS works (MI). Peptides (100%) were desalted using solid-phase extraction (SPE) with Wates pHLB Cl 8 plate. Peptides were loaded directly and eluted using 30 / 70 acetonitrile / water (0.1% TFA). Eluted peptides were lyophilized and reconstituted in 0.1% TFA. Peptides (50%) were analyzed in analytical duplicate by nano LC / MS / MS using a Waters NanoAcquity system interfaced to a ThermoFisher Fusion Lumos mass spectrometer. Peptides were loaded on a trapping column and eluted over a 75pm analytical column at 350nL / min; both columns were packed with Luna C18 resin (Phenomenex). A 2h gradient was employed The mass spectrometer was operated using a custom data-dependentDOCKET NO. STFD-011-PCT PROVISIONAL PATENT method, with MS performed in the Orbitrap at 60,000 FWHM resolution and sequential MS / MS performed using high resolution CID and EThcD in the Orbitrap at 15,000 FWHM resolution. All MS data were acquired from m / z 300-1600. A 3s cycle time was employed for all steps.Results
[0239] Single-chain pMHC trafficking to cell surface depends on specific peptide-HLA binding
[0240] The peptide-MHC single-chain trimer (pMHC SCT) has been extensively utilized to explore the interaction between pMHC and TCR. The pMHC consists of an 8-10 amino acid antigen peptide, beta2 microglobulin (B2M), and an MHC allele arranged in tandem and linked by flexible glycine and serine-rich linkers1'1,15. Typically, a high-affinity peptide for the MHC class I allele, preceded by a signal peptide (SP), is employed to traffic the trimeric complex to cell surface, endowing high cell surface expression in cells where endogenous MHC alleles were (HLA in humans) knocked-out (KO, FIGS 6A and 6B)
[0241] In our study, using the common HLA-A2 allele single-chain trimer, we observed that substituting the high-affinity peptide (HP) with one that lacked affinity for the HLA-A2 allele (no affinity, NP) led to a striking reduction of cell surface staining for the NP SCT. Neither antibodies targeting B2ni nor HLA- A2 showed positive staining ofNP construct (FIGS. 1A, IB, 6A, and 6B). To exclude misfolding as a potential cause for failed detection, we fused eGFP directly to the cytoplasmic tail of the HLA-A2 allele in both HP-SCT and NP-SCT. We observed robust membrane localization of the SCT with only the high-affinity pp65 peptide, whereas the trimer with a non-binding peptide displayed diffusive intracellular eGFP distribution (FIG. 1C), indicating that stable HLA136peptide binding is required for efficient surface trafficking of the pHLA complex. We further confirmed the structural integrity of surface-displayed SCTs via two independent approaches. First, HLA-A2:NY-ESO-1 SCT-expressing cells activated Jurkat cells expressing the cognate NY-ESO-1 -specific TCR evidenced by NF AT reporter gene induction, demonstrating proper peptide positioning for TCR recognition (FIG. 6C). Second, an antibody specific to H2-Kb:SIINFEKL (SEQ ID NO: 1303) showed concordant staining with total surface H2-Kb in SCT-transfected cells, indicating stable peptide binding (FIG. 6D). Together, these results showed that only properly folded SCTs with high-affinity peptides can traffic to andDOCKET NO. STFD-OH-PCT PROVISIONAL PATENT maintain structural integrity at the cell surface. Also, SCT with non-presentable peptides by HLA- A2 fails to traffic to the cell surface, resulting in low surface staining signal (FIGS. I A and IB).
[0242] To assess this phenomenon across multiple MHC alleles, we generated SCTs with representative alleles from HLA-A, B, and C groups (A*0101, B*0702, C*0401). Leveraging the IEDB training database, which contains peptides with measured binding affinity or from MS datasets for common HLA-I alleles, we fused peptides with high or no binding affinity systematically into SCTs for each HLA-I allele (Table 1). Consistently, trimers with peptides displaying no binding affinity exhibited markedly reduced surface expression across the examined HLA alleles (FIG. ID), a trend even observed with the mouse MHC H2-Kb allele (FIG. ID). The geometrical mean of fluorescent intensity (gMFI) for high-affinity peptides was - 1000-fold higher than that of no-binding-affinity peptides, offering a broad dynamic range for measurement (FIG. 6C and 6E). We observed a loose sigmoidal relationship between gMFI and peptide HLA-A2 affinities, reaching a plateau at higher affinities (FIG. IE, Table 2). The Y84C mutation conserved across HLA-A, B and C158 alleles and G2C in linker is commonly used to create a cysteinecysteine disulfide bridge! 59 and enhance the pHLA stability and surface display19-20(FIG. 6F). We compared pHLA surface display with or without G2C mutation across different peptide affinities and observed noisier surface straining without G2C mutation (FIG I E, Methods). This result suggests that G2C mutation enhances the sensitivity and robustness to detect HLA restricted peptides, largely due to the formation of disulfide trap from cysteine for the stabilization of pHLA complex10.
[0243] In the cellular context, HLA proteins are highly unstable without a bound peptide. The peptide-loading complex (PLC) stabilizes the MHC complex before a suitable peptide is loaded1'’18. We hypothesized that SCTs with non-presentable peptides are destabilized, impeding their exit from the ER and efficient trafficking to the cell membrane. Indeed, TAP1 / 2 knock-out minimally affected SCT presentation (FIGS. 6E and 6G), other than slightly lowering cell expression for all peptides irrespective of binding affinities. Further truncated SCTs with the peptide removed showed cytoplasmic diffusion, while HLA-A2 alone localized to the cell surface as expected in endogenous HLA KO cells (FIGS IF, 6D, 6E, 6H, and 61), supporting the notionDOCKET NO. STFD-011-PCT PROVISIONAL PATENT that SCT antigen presentation bypasses PLC. Additionally, when truncated SCTs were coexpressed with a presentable HLA peptide in the ER, surface presentation of the pHLA complex was partially restored (FIG IF, FIG. 61), albeit at least 10-fold less efficiently than the SCT. To a certain extent, this mirrors a prior method that used the co-expression of separate HLA and peptide transgenes to score antigen presentation13. These results suggest that cell surface pHLA expression can be used as a proxy for peptide binding strength and that SCTs offer a higher dynamic range and sensitivity.
[0244] ESCAPE-seq leverages HLA escape to cell surface to quantify peptide-HLA presentation
[0245] We next explored the potential of this system for high-throughput screening of HLA presented peptides (FIG. 2 A). Starting with peptides of known affinity from the IEDB training database, we selected thousands of peptides with different affinities in total for 4 representative HLA alleles (A*0101, A *0201 , B*0702 and C*0401). They were cloned into a pooled lentiviral vector per HLA allele, which was introduced into cells at low MOI (Methods). Cells were sorted into four intensity-based bins in log scale reflecting their ability of the programmed SCT to escape to cell surface (FIG. 2B). We assigned ESCAPE-scores (E-score) to each peptide (Methods), with higher E-score representing higher cell surface staining and presentation. Our biological replicates demonstrated high correlation (FIGS. 7 A and 7B) for all alleles studied using this method. Further, our E-score for the HLA-I allele generally correlated well with the measured IEDB affinity (c.g., Spearman coefficient of 0.79 for HLA-A*02, FIG 2C). We first compared the correlation between the E-score and IEDB affinity using either HLA-A*02 SCT harboring the Y84Cmutation (FIG. 2C) or the HLA-A*02 SCT without the mutation (FIG. 7C). Consistent with the results in Fig. IE, ESCAPE-seq on HLA-A*02 with the Y84C mutation shows a better correlation (Pearson correlation of 0.77 for HLA-A*02 SCT with Y84C vs.0.49 forHLA-A*02 without the mutation) Therefore, we used HLA-I alleles with the Y84Cmutation in SCT for all subsequent ESCAPE-seq experiments (FIG. 6F).
[0246] By binning IEDB affinity of defined presentable peptides into windows, we measured Recall and Precision rates (FIG. 2D, Methods). Comparison of ESCAPE-seq results with NetMHC4, a widely used pMHC prediction tool trained on IEDB 19, revealed ESCAPE’S reliableDOCKET NO. STFD-011-PCT PROVISIONAL PATENT performance. ESCAPE yielded >90% recall for high-affinity peptides, aligning closely with NetMHC4 predictions(FIG. 2D). The Receiver Operating Characteristics (ROC) curve indicated an area under curve (AUC) of 0.919 when evaluating ESCAPE against IEDB measured affinity, akin to NetMHC4 in both binding affinity (BA) and elution mode (EL) (FIG. 2E). We further compared with Precision-Recall curve (PRC), which offers a more robust metric when the positive rate i s low (typical for antigen presentation datasets where most pepti des will not bind the HLAs). Similarly, we observed comparable metrics in PRC, where ESCAPE obtains an AUC of 0.91 (FIG 2F).
[0247] Systematic comparisons across multiple HLA-I alleles including HLA-A*0101, B*0702, and C*0401 showed ESCAPE achieving similar performance to NetMHC (FIGS. 7D -7K), evident in both correlation coefficient (FIGS. 7L and 7M) and AUC-ROC (FIG. 2G). The PRC, a superior method for imbalanced data evaluation, mirrored the ROC findings (FIG. 21 1) Overall, our results indicate that ESCAPE-seq yields comparable metrics to NetMHC on the model's training data, which is the idealized scenario for NetMHC
[0248] ESCAPE-seq reveals presented epitopes from SARS-CoV-2 variants
[0249] Utilizing ESCAPE-seq, we next screened for presentable peptides in the Spike and N proteins of the SARS-CoV-2 virus. All peptides were standardized to a common length of 9 amino acids, comprising approximately 1500 peptides tiling across Spike and N protein, with an additional 900 viral mutation peptides from 17 SARS-CoV2 variant strains including alpha, beta, delta, and omicron (FIG. 3 A, Methods, Table 5).DOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0250] ESCAPE screens for SARS-CoV2 peptide pool on 3 common HLA-I alleles (A*0101, .A*0201, B*0702) were all consistent between replicates (FIGS. 8A and 8B). Comparing the E- scores from ESCAPE-seq against NetMHC predictions, we observed that - 90% of peptides were not presentable shown by both methods while a marginal ~ 2% of peptides are discordant between them, with slightly variation for different HLA-A alleles (FIG. 3B, FIGS. 8C---8E). The motif pattern of ESCAPE-positive peptides closely aligned with known patterns for HLA-I allele examples (FIG. 3C). Similarly, we identified that most literature-reported SARS-CoV-2 peptides appeared as positive by ESCA.PE-seq (FIGS 3D and 3E)20’21. Presentable peptides were dispersed across the viral proteins and seldom formed clusters or hot-spots in specific regions. Clustering peptides across A*01, A *02, and B*07 alleles indicated only a few' peptides were presented by two HLA alleles, and rarely by all three alleles (FIGS. 3F and 3G). We found that between 5-10% of peptides are generally presentable across 3 alleles depending on E-score cut-off (FIG. 8F). Interestingly, the AUC of the ROC between ESCAPE and NetMHC stood at approximately 0.95 when considering peptides from IEDB training pool, but decreased significantly for the SARS- CoV-2 peptides (FIGS, 8G and 8H). Since ESCAPE-seq performance is independent of peptide input pools, this decrease is likely due to decreased power of NetMHC prediction on new peptides.99 ]DOCKET NO. STFD-OH-PCT PROVISIONAL PATENTThis result suggests that NetMHC's performance is less effective on new peptides compared to its performance on familiar training data-an intuitively expected outcome.
[0251] Next we investigated changes in peptide presentation across multiple mutations in S and N protein from SARS-CoV-2 variant strains. Notably, numerous non-presentable peptides in the wild-type protein became presentable after introducing a mutation, and vice versa (FIGS. 3H, 31, 81, and 8J). Particularly interesting was the effect of point mutations, which altered the presentation status, resulting in a significantly lower percentage of peptides retaining their presentation status before and after a point mutation (FIG 3J, FIG. 8K).
[0252] ESCAPE-seq reveals landscape of cancer neoantigens from driver oncogenes presented by diverse HLA alleles, combinatorial ESCAPE-seq achieves simultaneous profiling of peptide presentation across diverse HLA alleles
[0253] We next integrated ESCAPE with a combinatorial HLA indexing strategy to screen peptide pools across multiple HLA alleles simultaneously (FIG. 4A, Methods). We integrated a barcode system with synonymous mutations in B2.M gene to enable HLA allele identification via sequencing (FIGS. 4A and 9 A). This multiplexed screening of peptide pools 261 with HLA allele pools can easily achieve >10, 000s of pHLA SCT variants in a single experiment (FIGS. 4A, 9B, and 9C Methods). The library' of SCT was transduced and sorted into four bins based on cell surface B2M expression as described above (FIGS. 2B, 9D, and 9E).
[0254] First, we conducted large-scale validation using existing MS data, leveraging a comprehensive monoalleiic HLA MS dataset2'. We selected top MS-captured peptides for 30 HLA alleles. We performed combinatorial ESCAPE-seq screening of these peptides across all 30 HLA alleles, examining over 29,000 pHLA interactions (FIG. 4H). Quality control analysis showed high consistency in both sequencing reads and calculated E-scores between replicates (FIGS 10A and 10B). As expected, MS270vali dated peptides had significantly higher E-score by ESCAPE-seq (FIG. 10C).
[0255] To assess ESCAPE-seq's ability to identify MS-validated peptides, we calculated the percentage of MS-detected peptides that were also identified by ESCAPE-seq foreach HLA allele (capture percentage, FIGS. 41). Remarkably, a single ESCAPE-seq experiment captured over 80%DOCKET NO. STFD-OH-PCT PROVISIONAL PATENT of MS-detected peptides across most HL A alleles, with some reaching 90%. This demonstrates that one ESCAPE-seq experiment can recapitulate the majority of findings from multiple MS experiments across 30 different HL A alleles. We further evaluated ESCAPE-seq's performance using receiver operating characteristic (ROC) curves, comparing true positive rates (TPR) against false positive rates (FPR) relative to MS data (FIGS. 41 and 10D). This analysis revealed strong performance, with AUC-ROC values between 0.8- 1.0 for most HLA alleles, although slightly lower values (—0.7) were observed for HLA-B56:01, B52:01, and C *07:01 (FIG. 4J). Additionally, when comparing with NetMHC4.1 predictions, we found ESCAPE-seq outperforms NetMHC4.1 for multiple HLA alleles (FIG. 10E).
[0256] Next, we performed ESCAPE-seq analysis with peptides ranging from 8 to 12 amino acids, examining 10 gene loci (3 mutations, 3 wild-type counterpart genes, 3 fusions. 1 deletion) across 25 HLA alleles (FIG. 10F). For each point mutation / fusion / deletion, we designed approximately 45 peptides of varying lengths that tiled across the mutation site or junction site, resulting in around 10,000 peptide-HLA interactions (FIGS. lOG and 10H). Among these peptides, 10-mers showed highest presentation frequency across examined alleles and targets, including KRASG12D which contains a well-established lOmer epitope (FIGS. 101 and 10J). An E293 score heatmap of mutation position versus peptide length shows that addition of one amino acid often maintains peptide presentation ability (FIG. 10J).DOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0257] Combinatorial ESCAPE-seq enables population-wide antigen presentation discovery of cancer neoantigens from driver oncogenes
[0258] To systematically identify cancer neoantigens across human populations by combinatorial ESCAPE-seq, we assembled 92 prevalent oncogenic mutations from cancer driver genes and 31 most frequent oncogenic fusions from the COSMIC database22(Methods). This curated pool encompassed point mutations such as Ras G12D, BRAF V600E. TP53 R175H and fusion junctions of oncogenic translocations such as BCR-ABL1 (Tables 6, 6.1, 6.2, and 6.3). We then created a peptide pool tiling over these point mutations and fusion junctions, and presented each peptide across 50 different HLA alleles (FIGS. 4A and 4K) Utilizing the 50 most common HL A -A, B, C alleles, accounting for over 90% of the global population29.
[0259] This multiplexed screening of peptide pools with HLA allele pools resulted in the generation of over 75,000 pHLA SCT variants in a single experiment (FIGS. 4A and 4K, Methods).DOCKET NO. STFD-011-PCT PROVISIONAL PATENTAdditionally, we spiked in 100 peptides of known antigens from the IEDB as controls. This extensive library produced a correlation coefficient of 0.89 between replicates (FIGS. 9A and 9B), indicating high reproducibility with scale.
[0260] We performed multiple optimizations to reduce mis-pairing between peptides and barcodes representing HLAs due to known recombination occurrences during pooled virus and DNA amplification24’23(.Methods). Further assessment, including a spike-in pool of peptides randomly paired with a subset of HLA. alleles, revealed potential recombination events during library' amplification and production After optimizations, we showed that these recombination-induced non-specific reads were low, less than 10% in our experiment (FIGS. 9C-9E).
[0261] The hi stogram of E-scores per HLA al lei e demonstrated two peaks with the dominant lower peak representing the negative peptide population (FIGS. 9E and 9F). Normalized E-scores were then calculated per peptide-HLA pair (Methods) For spike-in peptides selected from the IEDB, most exhibited a correlation between their E- scores and IEDB affinity or elution results (FIGS. 4B and 4L, Table 3). This consistency aligns with results obtained from IEDB benchmark pools for specific alleles (FIGS. 2A through 2H), indicating the robust performance of combinatorial ESCAPE-seq at the level of single HLA alleles.
[0262] Importantly, ESCAPE-seq uncovered previously reported tumor neoantigens paired with correct HLA alleles that were falsely predicted as non-binders by NetMHC (FIGS. 4C and 4M, Tables 7 and 7.1 ). This list is enriched of tumor neoantigens presented on HLA-C alleles such as KRAS G12V (GAVGVGKSA) / C*0304 (SEQ ID NO: 6) , KRAS G12C (GACGVGKSA) / C*0304 (SEQ ID NO: 7), KRAS G12D (GADGVGKSA) / C*0802 (SEQ ID NO: 8) and EGFR T790M (LTSTVQLIM) / C*070112’26’27(SEQ ID NO: 9), highlighting the sensitivity of ESCAPE-seq to identify HLA- C presented tumor neoantigens that are missed by computational prediction (FIGS. 4C and 4M). Notably, among 50 screened HLA alleles, we observed that reported HLA-C*0304 restricted KRAS G12V (GAVGVGKSA (SEQ ID NO: 10)) antigen can be presented on other HLA-I alleles to cover more diverse human population including B*5401 (East Asia), B*5601(Australia, Oceania), and C*0602 (North Africa)28 (Table 7).DOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0263] Table 7.1, below, lists the same tumor neoantigens as in Table 7 paired with other HLA al teles which were found to be binders via ESCAPE-seq. See FIGS. 4A-4M and 5A-5NDOCKET NO. STFD-011-PCT PROVISIONAL PATENTDOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0264] Table 8, below, List of peptides nominated by ESCAPE-seq used for validation assay forFIGS 16A-16I.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0265] The landscape of neoantigen presentation across oncogenes and HLA alleles (FIGS. 5 A and 5G) provided several lessons (FIG. 41)) Comparing HLA-A, B, and C alleles showcased significant variation in the percentage of presented peptides (FIGS. 4E, 9F, Methods). Notably, HLA-B alleles displayed substantial variance in peptide presentation among themselves, ranging from as low as 2% to exceeding 20%. Conversely, HLA-C alleles demonstrated a propensity to present a slightly higher number of peptides (FIGS, 4E and 9F), Most peptides are presented by just one or two HLA alleles (FIGS. 4F, 5H, 9G, and 9H), aligning with the evidence that different HLA alleles exhibited divergent motifs, length and amino acid preferences29'32. However, a small number of neoantigen peptides were presented by 20 or more HLA alleles (i.e., “public neoantigens”) (FIGS. 4F and 5H); such antigen peptides hold significant interest for potential utilization in immunotherapies. Using HLA-C *0304 as an example, the majority of these “public antigens” showed intermediate and high E-scores. Some of those with highest E- score could be readily detected and validated in independent MS experiments (FIGS. 51 and 9G, Methods).
[0266] Furthermore, focusing on functional differences among HLA-I alleles instead of merely their encoding sequences revealed a weak correlation between peptides presented by different HLA-I groups (FIGS. 9I-9K) For instance, a certain percentage of peptides could only be presented by HLA-As compared to HLA-Bs, while some were commonly presented by both HLA- A and B groups (FIG. 9N). Notably, peptides presented by both groups exhibited a weak correlation, suggesting that a peptide presented by multiple HLA-A alleles might have a higher likelihood of being presented by multiple HLA-B alleles, potentially due to its intrinsic propensity to bind in the peptide-binding groove. The functional similarity among HLA alleles can also be visualized by a uniform manifold approximation and projection (UMAP, FIG. 4G), showing the differences in antigen presentation in contrast to sequence-based clustering.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT
[0267] ESCAPE nominates high priority cancer neoantigens from driver oncoproteins
[0268] We next investigated antigen presentation characteristics of each oncogenic mutation by merging the results of all the peptides containing the same mutation across HLA alleles (FIG. 5D). Heatmap of the counts of presented peptides per mutation identified hot spots across HLAs (FIG. 5J). We define a presentable mutation as one with at least 1 presentable peptide tiling across this mutation. By this measure, many HLA alleles can present over 50% of mutations out of the 92 sampled here (FIGS 5B and 9H). In accordance with findings for individual peptides (FIGS. 4E and 9F), the mutation-centric measurement again confirmed the important role of HLA-C alleles in presenting oncogenic neoantigens, with the majority of the HLA-C alleles tested being able to present 60-80% of the mutations (FIGS. 5A, 5B, and 9H). By counting the HLAs that can present each mutation, we found that some mutations such as EGFR T790M andMED12 G44V can be presented by over 60% of the HLA alleles examined (“public” driver neoantigens), while other oncogenic mutations such as KILLS G12C and TP53 R282Ware only presented by ~2% alleles (“private” driver neoantigens, FIGS. 5 A, 5B, and 5N).
[0269] Considering that typical diploid human cells harbor two alleles each of HLA-A, B, and C, we extrapolated the distribution of presented mutation coverage within the human population through repeated sampling (Methods). Remarkably, a substantial majority (>90%) of the most prevalent oncogenic mutations could be presented by at least one allele, with a significant proportion of HLA sets covering 100% (FIG. 5C). By counting the HLAs that can present each mutation, we found that some mutations such as EGFR T790M and M YC S 16 IL can be presented by over 80% of the HLA alleles examined, while other oncogenic mutations such as SMAD4 R361H and TP53 R282W are only presented by -- 20% alleles (FIG. 5D).
[0270] Oncogenic fusion proteins are prevalent in many cancer types, in particular pediatric sarcomas33'34. Focusing on peptides derived from oncogenic fusion proteins, we observe similar trends in presentation rates across HLA alleles compared to peptides of wild-type or with mutations (FIG. 91). Aggregating peptides spanning the same fusion breakpoint highlights comparable variations in presentation coverage across HLA alleles (FIGS 5L and 9J), again confinning the ability of HLA-C alleles to broadly present oncogenic fusion neoantigens (FIGS 11C and 10D).DOCKET NO. STFD-011-PCT PROVISIONAL PATENTInterestingly, the distribution of personalized HLA presentation coverage of fusion mutations shifted left slightly less than the coverage of point mutations (FIG. HE), but the vast majority (>90%) are still covered and presented by the HL. As tested. We discovered a diverse range of antigen presentation compatibility for fusion oncoproteins (FIGS. 5E and 5M). For example, EWSR1-FLI fusion peptides are presented by almost 30 HLA alleles (-60% of HLAs tested, a "public neoantigen") whereas PAX7-FOXO fusion is only presented by one HLA. Of note, any presented peptides from fusion breakage points are considered as potential neoantigen as they are de novo peptides consisting of two gene fragments not observed in the wild type genome.
[0271] Considering that typical diploid human cells harbor two alleles of each HLA-A, B, and C, we extrapolated the distribution of presented mutation coverage within the human population through repeated sampling (Methods). Remarkably, a substantial majority (>90%) of the most prevalent oncogenic mutations could be presented by at least one allele (FIG. 9K). Interestingly, the distribution of personalized HLA presentation coverage of fusion mutations shifted left slightly less than the coverage of point mutations (FIG. 9K), but the majority (>80%) are still covered and presented by the HLAs tested.
[0272] Finally, we systematically assess the antigen presentation of oncogenic point mutations vs the wild type (WT) reference sequence. We expanded ESCAPE-seq library to include WT and oncogenic mutations so that they can be compared in the same assay. By comparing E-scores between mutant peptides and their wild-type counterparts, we identified four distinct outcomes: ( 1 ) neither mutant nor WT peptides are presented (WT394Mut-), (2) both mutant and WT peptides are highly presented on the cell surface, where the mutation may potentially impact TCR contact rather than HLA binding (WT+Mut+), (3)only WT peptides are presented, suggesting the mutation enables immune escape by preventing antigen presentation (WT+ only), and (4) only mutant peptides are presented, indicating the mutation confers novel HLA binding capability (Mut+ only) (FIG. 16A). It is known that a single amino acid change may not change the TCR binding or specificity33’36, thus differential antigen presentation is one of the key determinant therapeutic indexes of immunotherapies targeting cancer neoantigens. With HLA-A*0201 as an example, we found that approximately half of the peptides containing a point mutation and its correspondingDOCKET NO. STFD-011-PCT PROVISIONAL PATENT wild-type peptides are coordinately presented (FIG. 16B). Across all HLA alleles, the prevalence of this shared population where both mutant and WT equally presented is dominant (12A), aligning with the widely accepted notion that only a few amino acids on presented peptides serve as anchor points dictating binding37. Alternatively, we can display this quadrant in the perspective of per point408 mutation (FIG 16C, FIG. 12B), which highlights the peptides from one locus for all 50 alleles. Notable, a substantial number of mutant peptides are well presented but the corresponding WT peptide is not (FIGS. 16B, 16C, 12A, and 12B, mut + ). We reasoned that these mut+ peptides constitute promising immunogenic neoantigen candidates since immune tolerance of these antigens will not be established due to the lack of presentation of WT peptides. Indeed, we observed that these mut+ peptides contain previously reported immunogenic tumor neoantigens such as KRAS G12D (GADGVGKSA (SEQ ID NO: 8 )) / C*080226, FLT3 D835Y (YIMSDSNYV (SEQ ID NO: 599)) / A*020138, and EGFR T790M (IMQLMPFGC (SEQ ID NO: 459)) / A*020139 (FIGS. I6B and 16C). Thus, ESCAPE-seq permits nomination of potential immunogenic tumor neoantigens to prioritize candidates for antigen-directed immunotherapy.[0273| To test whether peptides that were selectively presented in their mutant form (Mut+only) are more immunogenic, we characterized T cell responses to EGFR T790Mneoantigen peptides in the context of HLA-A*23:01 , a prevalent FIFA allele in African American populations (FIGS. 16D and 16G, Methods). To characterize immunogenicity, we utilized a well -established in vitro peptide stimulation assay using PBMCs from an HLA-A*23:01+healthy donor. We differentiated antigen presenting cells (APCs) from donor PBMCs, stimulated them with peptide pools in the presence of adjuvants, and assessed T cell responses after 7 days of expansion46. Upon restimulation with individual peptides, w;e measured intracellular IFNy and TNFa production by flow cytometry' to determine T cell reactivity (FIG 16D, FIG. 12C). We compared T cell responses to three epitopes: one Mut+ peptide (pl 7: LTSTVQLIM (SEQ ID NO: 9)) and two Mut+WT+ peptides (p22:QLIMQLMPF (SEQ ID NO: 457) and p23: LIMQLMPFG (SEQ ID NO: 458)) (FIG. 16E, FIG. 12D). Flow' cytometry analysis revealed robust IFNv+ TNFa+ polyfunctional T cell responses exclusively to the Mut+ only peptide (FIGS. 16F and 16G). We then extended the T cell response assay to fusion on coproteins, which generate novel junctional sequences that couldDOCKET NO. STFD-011-PCT PROVISIONAL PATENT serve as immunogenic epitopes We selected 13 fusion breakpoint-derived epitopes 434 identified by ESCAPE-seq predicted to bind either HLA-A*03:01 or A*24:02 and tested them using PBMCs from an HLA-A*03:01 -r A*24:02-f- healthy donor. The analysis revealed that theNPM 1-ALK fusion-derived epitope was immunogenic, evidenced by a significant increase in polyfunctional T cells following peptide stimulation (FIGS. 1611, 161, and 12E). Together, these results establish ESCAPE-seq as a powerful tool for potentially immunogenic neoantigen discovery; with particular strength in identifying Mut+- only peptides that can bypass immune tolerance due to the lack of presentation of wild-type counterparts.Discussion
[0274] The HLA region is the most polymorphic segments within the human genome. With the rapid advance of sequencing technology, the number of identified HLA-A, B, and C gene alleles has surged, now numbering in over 26,000 within the current IMGT / HLA collection40. Developing a method capable of simultaneously' and efficiently screening antigen-presenting peptides across numerous HLA alleles could significantly accelerate our understanding of MHC-TCR biology. This includes the discovery of tumor neoantigens, the search for pathogenic antigens, and the mapping of novel TCRs, essential in advancing therapeutic interventions.((>275] MS-based HLA-I immunopeptidomes allow mapping of HLA-I eluted peptides and have been a mainstay of T cell antigen discovery. This approach is dominated by self peptides from the cellular proteome, limiting the sensitivity to detect clinically important peptides from pathogens and oncogenic proteins41. Further challenges for HLA-I antigen identification using MS include high cost, the requirement of specialized equipment, large quantities of sample material, ambiguous assignment of peptides to specific HLA allele, and potential peptide bias due to cysteine oxidation in MS sample preparation4,5,10,42. To address these limitations, DNA sequencing-based HLA immunopeptidomes discovery methods have emerged, leveraging large scale, cost-effective DNA oligo synthesis for antigen screening13. Instead of in vitro peptide binding to HLA, ESCAPE- seq and related methods13rely on the in vivo ER machinery as the quality control step for properDOCKET NO. STFD-011-PCT PROVISIONAL PATENT and stable pHLA folding as a prerequisite for cell surface trafficking. The single-chain design in ESCAPE-seq offers several ad...
Claims
DOCKET NO. STFD-011-PCT PROVISIONAL PATENTCLAIMS1. A composition comprising a first nucleic acid molecule comprising a first nucleic acid sequence comprising a first region, a second region, and a third region, wherein the first region, second region and third region collectively encode a single-chain polypeptide trimer, wherein the first region comprises a nucleic acid sequence encoding an antigen or antigenic determinant thereof, the second region comprises a nucleic acid sequence encoding a p2 microglobulin or a variant thereof, and the third region comprises a nucleic acid sequence encoding a first MHC class I allele or variant thereof; wherein the antigen or antigenic determinant thereof associates to the first MHC class I allele or variant thereof with an IC50 of no greater than about 500 nM.
2. The composition of claim 1 further comprising a second nucleic acid molecule comprising a second nucleic acid sequence that comprises a first region, a second region, and a third region, wherein the first, second and third regions collectively encode a single-chain polypeptide trimer; wherein the first region comprises a nucleic acid sequence encoding an antigen or antigenic determinant thereof, the second region comprises a nucleic acid sequence encoding a P2 microglobulin or a variant thereof, and the third region comprises a nucleic acid sequence encoding a second MHC class I allele or variant thereof; wherein the antigen or antigenic determinant thereof associates to the second MHC class I allele or variant thereof with an IC50 of no greater than about 500 nM.
3. The composition of claim 2, wherein the first MHC class I allele is selected from a first family of class I alleles, the second MHC class I allele is from a second family of class I alleles, and the first family of class I alleles is different than the second family of class I alleles.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT4 The composition of any of claims 2 or 3 further comprising a third nucleic acid molecule comprising a third nucleic acid sequence comprising a first region, a second region, and a third region, the first, second and third regions collectively encoding a single-chain polypeptide trimer, wherein the first region comprises a nucleic acid sequence encoding an antigen or antigenic determinant thereof, the second region comprises a nucleic acid sequence encoding P2microglobulin or a variant thereof, and the third region comprises a nucleic acid sequence encoding a third MHC class I allele or variant thereof; and wherein the antigen or antigenic determinant thereof associates to the third MHC class I allele or variant thereof with an IC50 of no greater than about 500 nM.
5. The composition of claim 4, wherein the first MHC class I allele is selected from a first family of class I alleles, the second MHC class I allele is from a second family of class I alleles, the third MHC class I allele is selected from a third family of class I alleles, wherein the first family of class I alleles is different than the second family of class 1 alleles, and the third family of class I alleles is different than the first family of class I alleles and the second family of class I alleles.
6. The composition of claims 4 or 5 further comprising a fourth nucleic acid molecule comprising a fourth nucleic acid sequence comprising a first region, a second region, and a third region, the first, second and third region collectively encoding a single-chain polypeptide trimer; wherein the first region comprises a nucleic acid sequence encoding an antigen or antigenic determinant thereof, the second region comprises a p2microglobulm or a variant thereof, and the third region comprises a fourth MHC class I allele or variant thereof; the antigen or antigenic determinant thereof associates to the fourth MHC class I allele or variant thereof with an IC50 of no greater than about 500 nM7. The composition of any one of claims 1 through 6, wherein the respective antigen or antigenic determinant thereof has an ESCORE from about 3.2 to about 5.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT8. The composition of claim 6 or 7, wherein the first MHC class I allele is selected from a first family of class I alleles, the second MHC class I allele is from a second family of class I alleles, the third MHC class I allele is selected from a third family of class I alleles, the fourth MHC class I allele is selected from a fourth family of class I alleles, the first family of class I alleles is different than the second family of class I alleles, the third family of class I alleles is different than the first family of class I alleles and the second family of class I alleles, and the fourth family of class I alleles is different that the first, the second, and the third families of class I alleles.
9. The composition of any of claims 4 through 7, wherein the first MHC class I allele is chosen from HLA-A alleles, and the second and the third MHC class I alleles are chosen from HLA-B or HLA-C alleles10 The composition of any of claims 3 through 7, wherein the first MHC class I allele is an allele chosen from HLA-A*01 alleles, and the second MHC class I allele is an allele chosen from HLA-A*02, HLA-A*03, HLA-A*024, HLA-A*026, HLA-B, and HLA-C alleles.
11. The composition of claim 6 or 7, wherein the first MHC class I allele is an allele chosen from HLA-A*01 alleles, the second MHC class I allele is an allele chosen from HLA- / V02, HLA- A*03, HLA-A*024, and HLA-A*026 alleles, the third MHC class I allele is chosen from HLA-B alleles, and the fourth MHC class I allele is chosen from HLA-C alleles.
12. The composition of any of claim 1 through 11, wherein at least the first nucleic acid molecule further comprises: (i) a nucleic acid sequence that encodes a flexible linker; or (ii) a first linker nucleic acid sequence and a second linker nucleic acid sequence, each of which encode a flexible linker or collectively encode one flexible linker; wherein, if the first nucleic acid molecule comprises the first and second linker nucleic acid sequence, the first linker nucleic acid sequence is positioned between and adjacent to the nucleicDOCKET NO. STFD-OH-PCT PROVISIONAL PATENT acid sequence the encodes an antigen or antigenic determinant thereof and the nucleic acid sequence that encodes the p2 microglobulin or variant thereof, and the second linker nucleic acid sequence is positioned between and adjacent to the nucleic acid sequence that encodes the |32 microglobulin or variant thereof and the nucleic acid sequence that encodes the MHC class I allele or variant thereof,13, The composition of any one of claims 1 through 12, wherein at least one of the nucleic acid molecules further comprises a nucleic acid sequence encoding a reporter.14, The composition of claim 12, wherein at least, the first nucleic acid sequence comprises a first linker nucleic acid sequence, a second linker nucleic acid sequence, and a third linker nucleic acid sequence, wherein the first, second and third linker each encode a flexible linker or the first, second or third linker nucleic acid sequences collectively encode a flexible linker, wherein the reporter nucleic acid sequence is positioned 3’ relative to the nucleic acid sequence encoding an MHC class I allele or variant thereof, and the third linker nucleic acid sequence is positioned between the nucleic acid sequence encoding the MHC class I allele or variant thereof and the reporter nucleic acid sequence15, The composition of any one of claims 1 through 14, wherein at least the first nucleic acid sequence further comprises a first 2A nucleic acid sequence encoding a first 2A sequence joined to a selection nucleic acid sequence encoding an antibiotic resistance protein or regulatory sequence.16, The composition of claim 15, wherein the respective first 2A nucleic acid sequence and respective selection nucleic acid sequence are positioned 3’ from the respective MHC nucleic acid sequence.DOCKET NO. STFD-OH-PCT PROVISIONAL PATENT17. The composition of claim 15, wherein at least one of the first nucleic acid sequence further comprises a second 2 A nucleic acid sequence encoding a second 2 A sequence joined to the reporter nucleic sequence and following the MHC nucleic acid sequence18, The composition of claim 17, wherein the second 2 A nucleic acid sequence joined to the reporter nucleic acid sequence is interposed between the MHC nucleic acid sequence and the first 2A nucleic acid sequence.
19. The composition of any one of claims 15 through 18, wherein the selection nucleic acid sequence encodes a puromycin N-acetyl -transferase.
20. The composition of any one of claims 15 through 19, wherein the reporter nucleic acid sequence encodes a GFP.
21. The composition of any of claims 1 through 20, wherein at least one of the first MHC nucleic acid sequence, the second MHC nucleic acid sequence, the third MHC nucleic acid sequence, and the fourth MHC nucleic acid sequence encodes an HLA class I allele or variant thereof.
22. The composition of any one of claims 1 through 21, wherein at least one of the first nucleic acid sequence, the second nucleic acid sequence, the third nucleic acid sequence, and the fourth nucleic acid sequence further comprise a respective vector backbone.23, The composition of claim 22, wherein at least one of the respective vector backbone nucleic acid sequence is a lentivector sequence.
24. The composition of any one of claims 1 through 23, wherein the first MHC class I allele is an HLA class I allele chosen from one or a combination of two or more of A*01 :01, A*02:01,DOCKET NO. STFD-011-PCT PROVISIONAL PATENTA *02'05, A*02: 12, A*03 :01 , A* 11 :01 , A*23 :01 , A *24:02, A*30:()l, A*31 :01 , A*31 :08, A*34:01 , A*33:03, A*68:01, B*07:02, B*08:01, B*08:02, B*13:01, B*15:01, B*15:02, B* 15:03, B*18:01, 8*27:01 , B*27:05, B*27:02, B* 35:01, B*35:02, B*35:08, B*39:06, B*40:0L B *40:06, B*40: 10, B*41 :01, B*44:0, B*46:01, B*48:03, B*50:01, B*51 :01, B*52:01, B*54:01, B*56:01, B*57:01, B*58:01, C*03:04, C*04:01, C*06:02, C*07:01, C*07:02, C*08:02, C* 12:03.
25. The composition of any one of claim 1 through 24, wherein the second MHC class I allele is an HLA class I allele chosen from one or a combination of two or more of: A*01:01, A*02:01, 4*02'05. A*02: 12, A*03:01, A* 11 :01, A*23:01, A*24:02, A*30:01, A*31 :01, A*31 :08, A*34:01, A*33:03, A*68:01, B*07:02, B*08:01, B*08:02, B*13:01, B* 15:01 , B*15:02, B*15:03, B*18:01 , B*27:
01. B*27:05, B*27:02, B*35:O1, B*35:02, B*35:08, B*39:06, B*40:01, B*40:06, B*40: 10, 8*41 :01 , B*44:0, B*46:01, B*48;03, 8*50:01, B*51 :01 , B*52:01 , B*54:01, B* 56:01, B*57:01, 8*58:01, C*03:04, C*04:01, C*06:02, C*07:01, C*07:02, C*08:02, C* 12:03.
26. The composition of any one of claims 1 through 25, wherein the third MHC class I allele is an HLA class I allele chosen from one or a combination of two or more of: A*01 :01, A*02:01, A*02:05, A *02: 12, A*03 :01 , A* 11 :01 , A *23 :01 , A*24:02, A*30:01 , A*31 :01 , A*31 :08, A*34:01 , A*33:O3, A*68:01, B*07:02, B*08:01, B*08:02, B*13:01, B* 15:01, B*15:02, B*15:O3, B* 18:01, B*27:01, B*27:05, B*27:02, B*35:01, B*35:02, B*35:08, B*39:06, B*40:01, B*40:06, B*40: 10, B*41 :01, B*44:0, B*46:01, B*48:03, B*50:01, B*51 :01, B*52:01, B*54:01, B*56:01, B*57:01, 8*58:01, C*03:04, C*04:01, C*06:02, C*07:01, C*07:02, C*08:02, C*12:03.
27. The composition of any one of claims 1 through 26, wherein the fourth MHC class I allele is an HLA class I allele chosen from one or a combination of two or more of: A*01:01, A*02:01 , .V 02.05, A*02: 12, A*03:01, A* 11 :01, A*23:01, A*24:02, A*30:01, A*31 :01, A*31 :08, A*34:01, A*33:03, A*68:01, B*07:02, 8*08:01, B*08:02, B*13:01, B*15:01, B*15:02, B *15:03, B*18:01, B*27:01, B*27:05, B*27:02, B*35:
01. B*35:02, B*35:08, B*39:06, B*40:01, B*40:06, B*40: 10,DOCKET NO. STFD-011-PCT PROVISIONAL PATENTB*41 :01, B*44:0, 8*46:01, 8*48:03, B*50:01, B*51 :()1, B*52:01, B*54:01, 8*56:01 , 8*57:01 , B*58:01, C*03:04, C*04:01, C*06:02, C*07:01, C*07:02, C*08:02, C* 12:03.
28. A composition comprising: a first nucleic acid molecule comprising a first nucleic acid sequence encoding a first single-chain polypeptide trimer comprising a first region, a second region, and a third region, the first region comprising an antigen or antigenic determinant thereof, the second region comprising p2microglobulin or a variant thereof, and the third region comprising a first MHC class I allele or variant thereof, wherein the antigen or antigenic determinant thereof is encoded by an antigen nucleic acid sequence within the first nucleic acid sequence, the p2microglobulin or a variant thereof is encoded by a |32 nucleic acid sequence within the first nucleic acid sequence, and the first MHC class I allele or variant thereof is encoded by an MHC nucleic acid sequence within the first nucleic acid sequence, and a second nucleic acid molecule comprising a second nucleic acid sequence encoding a second single-chain polypeptide trimer comprising a first region, a second region, and a third region, the first region comprising an antigen or antigenic determinant thereof, the second region comprising p2microglobulin or a variant thereof, and the third region comprising a second MHC class I allele or variant thereof, wherein the antigen or antigenic determinant thereof is encoded by an antigen nucleic acid sequence within the second nucleic acid sequence, the p2microglobulin or a variant thereof is encoded by a p2 nucleic acid sequence within the second nucleic acid sequence, and the second MHC class I allele or variant thereof is encoded by an MHC nucleic acid sequence within the second nucleic acid sequence, wherein the first single-chain polypeptide trimer is different than the second single-chain polypeptide trimer.
29. The composition of claim 28 further comprising: a third nucleic acid molecule comprising a third nucleic acid sequence encoding a third single-chain polypeptide trimer comprising a first region, a second region, and a third region, theDOCKET NO. STFD-011-PCT PROVISIONAL PATENT first region comprising an antigen or antigenic determinant thereof, the second region comprising P2microgiobulin or a variant thereof, and the third region comprising a first MHC class I allele or variant thereof, wherein the antigen or antigenic determinant thereof is encoded by an antigen nucleic acid sequence within the third nucleic acid sequence, the P2microglobulin or a variant thereof is encoded by a p2 nucleic acid sequence within the third nucleic acid sequence, and the third MHC class I allele or variant thereof is encoded by an MHC nucleic acid sequence within the third nucleic acid sequence, wherein the first single-chain polypeptide trimer is different than the second single-chain polypeptide trimer, and the third single-chain polypeptide trimer is different than the first and the second single-chain polypeptide trimers.
30. The composition of claim 29 further comprising: a fourth nucleic acid molecule comprising a fourth nucleic acid sequence encoding a fourth single-chain polypeptide trimer comprising a first region, a second region, and a third region, the first region comprising an antigen or antigenic determinant thereof, the second region comprising P2microglobulin or a variant thereof, and the third region comprising a second MHC class I allele or variant thereof, wherein the antigen or antigenic determinant thereof is encoded by an antigen nucleic acid sequence within the fourth nucleic acid sequence, the p2microglobulin or a variant thereof is encoded by a p2 nucleic acid sequence within the fourth nucleic acid sequence, and the fourth MHC class I allele or variant thereof is encoded by an MHC nucleic acid sequence within the fourth nucleic acid sequence, wherein the first single-chain polypeptide trimer is different than the second single-chain polypeptide trimer, the third single-chain polypeptide trimer is different than the first and the second single-chain polypeptide trimers, and the fourth single-chain polypeptide trimer is different than the first, second, and third single-chain polypeptide trimers.
31. The composition of any one of claims 28 through 30, wherein the respective antigen or antigenic determinant thereof has an ESCORE from about 3.2 to about 5.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT32. The composition of claim 30 or 31, wherein the first MHC class I allele is selected from a first family of class I alleles, the second MHC class I allele is from a second family of class I alleles, the third MHC class I allele is selected from a third family of class I alleles, the fourth MHC class I allele is selected from a fourth family of class I alleles, the first family of class I alleles is different than the second family of class I alleles, the third family of class I alleles is different than the first family of class I alleles and the second family of class I alleles, and the fourth family of class I alleles is different that the first, the second, and the third families of class I alleles.
33. The composition of any of claims 28 through 31, wherein the first MHC class I allele is chosen from HLA-A alleles, and the second and the third MHC class I alleles are chosen from HLA-B or HLA-C alleles.
34. The composition of any of claims 28 through 31, wherein the first MHC class I allele is chosen from HLA-A*01 alleles, and the second MHC class I allele is chosen from HLA-A*02, HLA-A*03, HLA-A*024, HLA-A*026, HLA-B, and HLA-C alleles.
35. The composition of claim 30 or 31, wherein the first MHC class I allele is chosen from HLA-A*01 alleles, the second MHC class I allele is chosen from HLA-A*02, HLA-A*03, HLA- A*024, and HLA-A* 026 alleles, the third MHC class I allele is chosen from HLA-B alleles; and the fourth MHC class I allele is chosen from HLA-C alleles.
36. The composition of any of claim 28 through 35, wherein at least one of the first nucleic acid molecule, the second nucleic acid molecule, the third nucleic acid molecule, or the fourth nucleic acid molecule further comprises a first flexible nucleic acid sequence and second flexible nucleic acid sequence that together encode a flexible linker, wherein the first flexible nucleic acid sequence is positioned between and adjacent to the respective antigen nucleic acid sequence1andDOCKET NO. STFD-011-PCT PROVISIONAL PATENT the respective p2 nucleic acid sequence, and the second flexible nucleic acid sequence is positioned between and adjacent to the 3’ end of the respective (32 nucleic acid sequence and the respective MHC nucleic acid sequence.
37. The composition of any one of claims 28 through 36, wherein at least one of the first nucleic acid sequence, the second nucleic acid sequence, the third nucleic acid sequence, and the fourth nucleic acid sequence further comprise a reporter nucleic acid sequence encoding a reporter38. The composition of claim 37 further comprising a third flexible nucleic acid sequence encoding a third flexible linker, wherein the reporter nucleic acid sequence is positioned 3’ relative to the MHC nucleic acid sequence, and the third flexible nucleic acid sequence between and joining the MHC nucleic acid sequence and the reporter nucleic acid sequence.
39. The composition of any one of claims 28 through 38, wherein at least one of the first nucleic acid sequence, the second nucleic acid sequence, the third nucleic acid sequence, and the fourth nucleic acid sequence further comprise a first 2A nucleic acid sequence encoding a first 2A sequence joined to a selection nucleic acid sequence encoding an antibiotic resistance protein or regulatory sequence.
40. The composition of claim 39, wherein the respective first 2A nucleic acid sequence and respective selection nucleic acid sequence are positioned 3’ from the respective MHC nucleic acid sequence.
41. The composition of claim 39, wherein at least one of the first nucleic acid sequence, the second nucleic acid sequence, the third nucleic acid sequence, and the fourth nucleic acid sequence further comprise a. second 2A nucleic acid sequence encoding a second 2A sequence joined to the reporter nucleic sequence and following the MHC nucleic acid sequence.DOCKET NO. STFD-OH-PCT PROVISIONAL PATENT42. The composition of claim 41 , wherein the second 2A nucleic acid sequence joined to the reporter nucleic acid sequence is interposed between the MHC nucleic acid sequence and the first 2A nucleic acid sequence.
43. The composition of any one of claims 39 through 42, wherein the selection nucleic acid sequence encodes a puromycin N-acetyl-transferase.
44. The composition of any one of claims 39 through 43, wherein the reporter nucleic acid sequence encodes a GFP45. The composition of any of claims 28 through 44, wherein at least one of the first MHC nucleic acid sequence, the second MHC nucleic acid sequence, the third MHC nucleic acid sequence, and the fourth MHC nucleic acid sequence encodes an HLA class I allele or variant thereof.
46. The composition of any one of claims 28 through 45, wherein at least one of the first nucleic acid sequence, the second nucleic acid sequence, the third nucleic acid sequence, and the fourth nucleic acid sequence further comprise a respective vector backbone nucleic acid sequence.
47. The composition of claim 46, wherein at least one of the respective vector backbone nucleic acid sequence is a lentivector sequence.
48. The composition of any one of claims 28 through 47, wherein the first MHC class I allele is an HLA class I allele chosen from one or a combination of two or more of: A*01:01, A*02:01, .V 02.05, A*02: 12, A*03:01, A* 11 :01, A*23:01, A*24:02, A*30:01, A*31 :01, A*31 :08, A*34:01, A*33:03, A*68:01 , B*07:02, 8*08:01, B*08:02, B*13 :() 1 , B*15:01, 6*15:02, B *15:03, B*18:01, B*27:01, B*27:05, B*27:02, B*35:
01. B*35:02, B*35:08, B*39:06, B*40:01, B*40:06, B*40: 10,DOCKET NO. STFD-OH-PCT PROVISIONAL PATENTB*41 :01, B*44:0, 8*46:01, 8*48:03, B*50:01, B*51 :()1, B*52:01, B*54:01, 8*56:01 , 8*57:01 , B*58:01, C*03:04, C*04:01, C*06:02, C*07:01, C*07:02, C*08:02, C* 12:03.
49. The composition of any one of claims 28 through 48, wherein the second MI IC class I allele is an HLA class I allele chosen from one or a combination of two or more of. A*01:01, A*02:01, A*02:05, A*02: 12, A*03:0l, A*H :01, A*23:01, A*24:02, A*30:01, A*31 :01, A*31:08, A*34:0l, A*33:03, A*68:01, B*07:02, B*08:01, B*08:02, B*13:01, B*15:01, B*15:02, B*15:03, B*18:01, 8*27:01, B*27:
05. 8*27:02, B*35:01, 8*35:02, B*35:08, 8*39:06, B*40:01 , B*40:06, B*40:10, B*41:01, B*44:0, B*46:01, B*48:03, B*50:01, B*5I:01, B!*52:01, B*54:01, B*56:01, B*57:01, B*58:01, C*03:04, C*04:01 , C*06:02, C*07:01, C*07:02, C*08:02, C*12:03.
50. The composition of any one of claims 28 through 49, wherein the third MHC class I allele is an HLA class I allele chosen from one or a combination of two or more of: A*01 :01, A*02:01, A*02:05, A*02:12, A*03:01, A* 11 :01, A*23;0L A*24:02, A*30:01, A*31 :01 , A*31:08, A.*34:0l, A*33:03, A*68:01, 8*07:02, B*08:01, B*08:02, 8*13:01, B*15:01, 8*15:02, B*15:03, 8*18:01, B*27:01, 8*27:05, B*27:02, 8*35:01, B*35:02, B*35:08, B*39:06, B*40:01, B*40:06, B*40:10, 8*41 :01 , 8*44:0, B*46:01, B*48:03, 8*50:01 , 8*51 :01 , 8*52:01 , B*54:01 , B*56:01, B*57:01 , B*58:01, C*03:04, C*04:01, ( 1*06:02, C*07:01, C*07:02, ('*08:
02. C* 12:03.
51. The composition of any one of claims 28 through 50, wherein the fourth MHC class I allele is an HLA class I allele chosen from one or a combination of two or more of: A*01 :01, A*02:01, A*02:05, A*02: 12, A*03:01, A* I 1 :01, A*23:01 , A*24:02, A*30:01, A*31 :01, A*31 :08, A*34:01, A*33:03, A*68:01, 8*07:02, B*08.01, B*08:02, 8*13:01, B* 15:01, B*15:02, B*15:03, B* 18:01 , B*27:01, 8*27:05, B*27:02, 8*35:01, B*35:02, 8*35:08, B*39:06, 8*40:01, B*40:06, 8*40: 10, B*41 :01, B*44:0, B*46:01, B*48:03, B*50:01, B*51 :01, B*52:01, B*54:01, B*56:01, B*57:01, 8*58:01, C*03:04, C*04:01, C*06:02, C*07:01, C*07:02, C*08:02, C* 12:03.25DOCKET NO. STFD-011-PCT PROVISIONAL PATENT52. The composition of any one of claims 1 through 51 lacking each antigen or antigenic determinant thereof and each antigen nucleic acid sequence is replaced by a respective insertion site.
53. The composition of claim 52 further comprising a respective peptide encoding sequence within the respective insertion site.
54. The composition of claim 53, wherein the peptide encoding sequence encodes a peptide of 8—10 amino acids.
55. The composition of any one of claims 1 through 51, wherein at least one of the first nucleic acid molecule, the second nucleic acid molecule, the third nucleic acid molecule, and the fourth nucleic acid molecule further comprise a signal nucleic acid sequence encoding a signal peptide; wherein the respective nucleic acid molecule comprises in 5’ to 3’ orientation the signal nucleic acid sequence, the antigen nucleic acid sequence, the p2 nucleic acid sequence, the MHC nucleic acid sequence.
56. A cell comprising the composition of any of claims 1 through 55.
57. A cell comprising the composition of any one of claims 1 through 51.
58. A kit comprising the composition of any of claims 1 through 55.
59. The kit of claim 58 comprising a cell capable of being transformed with one or more of the first nucleic acid molecule, the second nucleic acid molecule, the third nucleic acid molecule, and the fourth nucleic acid molecule.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT60. A method of identifying an immunotherapy target comprising selecting from a population of cells comprising the composition of any one of claims 1 through 51 a subpopulation of cells surface displaying the respective single-chain polypeptide trimer, and identifying the respective antigen or antigenic determinant thereof within the respective single-chain polypeptide trimer for at least one cell in the subpopulation.
61. The method of claim 60, wherein the identifying comprises identifying the respective antigen or antigenic determinant thereof within the respective single-chain polypeptide trimer for each cell in the subpopulation.
62. The method of claim 60 or 61 further comprising calculating an E-score for each identified antigen or antigenic determinant.63 The method of claim 62 selecting an identified antigen or antigenic determinant thereof with a high E-score as the immunotherapy target.
64. The method of claim 63, wherein the high E-score is about 3.2 to about 5.
65. The method of claim 63 or 64, wherein calculating the E-score comprises normalizing all reads within each bin by dividing the counts in each bin by the average, normalizing the total number of sequencing reads measured per trimer (pl ILA pair), and synthesizing these data into a singular E-score.
66. The method of claim 65, wherein the step of synthesizing comprises application of the following formula: Escore = countsjag * vv_bg + counts low *w_low + coimts med * w_med+coimts_high* w_high where w_bg = 0,w_low = 2,w_med =4,tv_high= 8, where weight is at log scale matching the binning scale during cell sorting, while the background bin is assigned 0.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT67. The method of claim 60 or 61, wherein an identified antigen or antigenic determinant is the immunotherapy target68, The method of any one of claims 60 through 67, wherein the identifying comprises sequencing the antigen nucleic acid sequence.
69. The method of any one of claims 60 through 67 further comprising transfecting cells with the composition of any one of claims 1 through 51 to create the population of cells.
70. The method of claim 69 further comprising cloning an antigen nucleic acid sequence into the insertion site of the composition of claim 52.71 The method of any one of claims 60 through 67 further comprising transfecting cells with a nucleic acid molecule library, wherein a plurality of nucleic acid molecules in the nucleic acid library comprise a nucleic acid sequence encoding a single-chain polypeptide trimer comprising a first region, a second region, and a third region, the first region comprising an antigen or antigenic determinant thereof, the second region comprising plmicroglobulin or a variant thereof, and the third region comprising an MHC class I allele or variant thereof, wherein the antigen or antigenic determinant thereof is encoded by an antigen nucleic acid sequence within the nucleic acid sequence, the p2microglobulin or a variant thereof is encoded by a P2 nucleic acid sequence within the nucleic acid sequence, and the first MHC class I allele or variant thereof is encoded by an MHC nucleic acid sequence within the nucleic acid sequence, each of the antigen nucleic acid sequence, the P2 nucleic acid sequence, and the MHC nucleic acid sequence comprising a respective 5’ end and respective 3’ end, and two or more nucleic acid molecules in the nucleic acid library differ from each other in at least one of the antigen nucleic acid sequence and the MHC nucleic acid sequence.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT72. The method of claim 71, wherein the MHC nucleic acid sequence in different ones of the two or more nucleic acid molecules encodes and HLA-A, HLA-B, or HLA-C allele.
73. The method of claim 72, wherein the HLA-.A, HLA-B, or HLA-C allele are independently and respectively selected from A*01:01, A*02:01, A*02:05, A*02:12, A*03:01, A* 11 :01 , A*23:0l, A*24:02, A*30:01, A*31 :01, A*31 :08, A*34:01, A*33:03, A*68:01, B*07:02, B*08:01, B *08:02, B* 13:01, B*15:01 , B*15:02, B*15:03, B*18:01, B *27:01, B*27:05, B*27:02, B*35:0L B*35:02, B*35:08, B*39:06, B*40:
01. B*40:06, B*40: 10, B*4l:01, B*44:0, B*46:01, B*48:03, B*50:01, B*51:01, B*52:01, B*54:01, B*56:01, B*57:01 , B*58:01, C*03:04, C*04:01, C*06:02, C;07.
01. C*07:02, C*08:02, or C* 12:03.
74. The method of claim 73, wherein a set of the different ones of the two or more nucleic acid molecules encodes each of A*01:01, A*02:01, A*02:05, A*02:12, A*03:01, A*ll :01, A*23:01, A *24:02, A*30:01, A*31 :01, A*31 :08, A*34:01, A*33:03, A*68:01, B*07:02, B;08.01 , B*08:02, B*13:01, B*15:01, B*15:02, B*15:O3, B*18:01, B*27:01, B*27:05, B*27:02, B*35:01, B*35:02, B*35:08, B*39:06, B*40:01, B*40:06, B*40: 10, B*41:01, B*44:0, B*46:01, B*48:03, B*50:01, B*51 :01 , B*52:01 , B*54:01 , B*56:01, B*57:01 , B*58:01, C*03:04, ('*04.0 ! . C*06:02, ('*07:01 . C*07:02, C*08:02, or C* 12:03 on separate ones of the two or more nucleic acid molecules.
75. The method of claim 72, wherein a set of the different ones of the two or more nucleic acid molecules encodes two or more of an HLA-A allele, an HLA-B allele, or HLA-C allele.
76. The method of any of claim 71 through 75, wherein the nucleic acid molecule further comprises a first flexible nucleic acid sequence and second flexible nucleic acid sequence that together encode a flexible linker, wherein the first flexible nucleic acid sequence is positioned between and adjacent to the antigen nucleic acid sequence and the |32 nucleic acid sequence, and the second flexible nucleic acid sequence is positioned between and adjacent to the 3’ end of the [32 nucleic acid sequence and the MHC nucleic acid sequence.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT77. The method of any one of claims 71 through 76, wherein the nucleic acid sequence further comprises a reporter nucleic acid sequence encoding a reporter78, The method of claim 77, wherein the nucleic acid molecule further comprises a third flexible nucleic acid sequence encoding a third flexible linker, wherein the reporter nucleic acid sequence is positioned 3’ relative to the MHC nucleic acid sequence, and the third flexible nucleic acid sequence between and joining the MHC nucleic acid sequence and the reporter nucleic acid sequence.
79. The method of any one of claims 71 through 78, wherein the nucleic acid sequence further comprises a first 2A nucleic acid sequence encoding a first 2A sequence joined to a selection nucleic acid sequence encoding an antibiotic resistance protein or regulatory sequence.
80. The method of claim 79, wherein the first 2A nucleic acid sequence and selection nucleic acid sequence are positioned 3’ from the MHC nucleic acid sequence.
81. The method of claim 80, wherein the nucleic acid sequence further comprises a second 2A nucleic acid sequence encoding a second 2 A sequence joined to the reporter nucleic sequence and following the MHC nucleic acid sequence.
82. The method of claim 81 , wherein the second 2A nucleic acid sequence joined to the reporter nucleic acid sequence is interposed between the MHC nucleic acid sequence and the first 2A nucleic acid sequence.
83. The method of any one of claims 79 through 82, wherein the selection nucleic acid sequence encodes a puromycin N-acetyl-transferase.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT84. The method of any one of claims 77 through 83, wherein the reporter nucleic acid sequence encodes a GFP.
85. The method of any of claims 71 through 84, wherein the MHC nucleic acid sequence encodes an HLA class I allele or’ variant thereof86. The method of any one of claims 71 through 85, wherein the nucleic acid sequence further comprise a vector backbone nucleic acid sequence.
87. The method of claim 86, wherein the vector backbone nucleic acid sequence is alentivector sequence.
88. The method of any one of claims 60 through 87, wherein selecting comprises detecting cells with surface expressed j32microglobulin and / or MHC class I allele.
89. The method of any one of claims 71 through 88 further comprising inserting a plurality of antigen nucleic acid sequences into an insertion site of a nucleic acid cassette comprising the p2 nucleic acid sequence within the nucleic acid sequence and the MHC nucleic acid sequence to create the nucleic acid molecule.
90. The method of claim any one of claims 60 through 89, wherein the antigen nucleic acid sequence is selected from coding sequences of one or more of a viral protein, an oncoprotein, or a non-viral intracellular pathogen protein,91. The method of any one of claims 60 through 90, wherein the cells are HLA / TAP knockout cell s.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT92. The method of any of claims 60-91 , wherein in the antigen or antigenic determinant thereof is from about 8 to about 10 amino acids in length.
93. A system comprising the composition of any of claims 1-55 and HL / VTAP knockout cells.
94. A method of making the composition of any one of claims 1 through 51 comprising cloning one or more antigen sequence into the insertion site of the composition of claim 52.
95. The method of claim 94, wherein the one or more antigen sequence comprises a library of antigen sequences.
96. A composition comprising any antigen or antigenic determinant thereof disclosed herein.97 A pharmaceutical composition comprising any antigen or antigenic determinant thereof disclosed herein and a pharmaceutically acceptable carrier.
98. A method of treating a subject comprising administering a therapeutically effective amount of the composition of claim 96 or the pharmaceutical composition of claim 97 to the subject.
99. A dual HL A and TAP knock-out cell.
100. A composition comprising a dual HLA and TAP knock-out cell.
101. The composition of claim 100, wherein the cell comprises any one or more nucleic acid disclosed herein.
102. The composition of claim 100 or 101, wherein the cell comprises any one or more amino acid sequence disclosed herein.DOCKET NO. STFD-011-PCT PROVISIONAL PATENT103. A composition comprising any one or more amino acid sequence disclosed herein directly or translated from nucleic acid sequence.104 A pharmaceutical composition comprising any one or more amino acid sequence disclosed herein directly or translated from nucleic acid sequence and a pharmaceutically acceptable carrier.
105. An engineered T-cell receptor comprising any one or more antigen or antigenic determinant herein.
106. A nucleic acid molecule comprising a nucleic acid sequence encoding the engineered T- cell receptor of claim 105.107 An engineered T-cell comprising the engineered T-cell receptor of claim 105 or the nucleic acid molecule of claim 106.
108. A nucleic acid molecule comprising a nucleic acid sequence encoding a plurality of antigen or antigenic determinants herein.
109. A vector comprising a nucleic acid molecule including a nucleic acid sequence encoding any one or more antigen or antigenic determinant herein.110 A cell comprising the engineered T-cell receptor of claim 105, the nucleic acid molecule of claim 106, the nucleic acid molecule of claim 108, or the vector of claim 108.
111. A pharmaceutical composition comprising the engineered T-cell receptor of claim 105, the nucleic acid molecule of claim 106, the engineered T-cell of claim 107, the nucleic acid moleculeDOCKET NO. STFD-011-PCT PROVISIONAL PATENT of claim 108, the vector of claim 109, or the cell of claim 110 and a pharmaceutically acceptable carrier.
112. A vaccine comprising the engineered T-cell receptor of claim 105, the nucleic acid molecule of claim 106, the engineered T-cell of claim 107, the nucleic acid molecule of claim 108, the vector of claim 109, the cell of claim 110, or the pharmaceutically composition of claim 111.
Citation Information
Patent Citations
Personalized cancer vaccines and methods therefor
US20170202939A1
Single chain trimer MHC class i nucleic acids and proteins and methods of use
US20240239869A1
Methods of screening for peptide-HLA class i alloreactivity
WO2023239736A1