Novel protein isoforms and uses

A method identifies and utilizes novel protein isoforms from non-canonical splicing, especially those with transposable elements, for targeted cancer therapy and diagnosis, improving therapeutic and diagnostic outcomes.

WO2026062222A1PCT designated stage Publication Date: 2026-03-26INSTITUT CURIE +2
View PDF 62 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2026-03-26

Smart Images

  • Figure EP2025076896_26032026_PF_FP_ABST
    Figure EP2025076896_26032026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides novel protein isoforms, including transmembrane proteins, encoded by variant transcripts generated from non-canonical splicing, and corresponding peptides, nucleic acids, vaccines, antibodies and immune cells that can be used in diagnosis and therapy.
Need to check novelty before this filing date? Find Prior Art

Description

NOVEL PROTEIN ISOFORMS AND USESFIELD

[0001] The present disclosure provides novel protein isoforms, including transmembrane proteins, generated from non-canonical splicing, and corresponding peptides, nucleic acids, vaccines, antibodies and immune cells that can be used in diagnosis and therapy.BACKGROUND

[0002] Different mechanisms can generate non-canonical proteins in mammals. Novel open reading frames (ORFs) from non-genic regions can encode microproteins or short ORFs. Microproteins are evolutionary young amino acid sequences that can be involved in development, metabolism, or cancer, for example. Novel ORFs can also overlap existing protein coding regions, but in a different reading frame, a process called “overprinting.” New variants of existing proteins are generally generated through either mRNA splicing or gene duplication.

[0003] During evolution, mRNA splicing allows the exonization of non-coding regions, generally from introns, generating evolutionary recent exons. Cryptic splice motifs emerge in introns or in vicinity of genes through random mutations, providing opportunities for the splicing machinery to generate novel non-canonical spliced transcripts. If evolutionarily advantageous, such new splicing sites can increase in frequency in the population through positive selection and become a source of alternative, and even constitutive, spliced isoforms. A significant proportion of these emerging splice signals are located within intronic transposable elements (TEs).

[0004] Mammalian TEs are divided in four main classes: DNA transposons and three classes of retrotransposons: short interspersed nuclear elements (SINEs), long interspersed nuclear elements (LINEs), and long terminal repeats (LTRs). Each TE class is subdivided into families and subfamilies (over 1000 sub-families in humans). De novo TE exonization has been studied in large collections of cancer and healthy tissue samples, based on the detection of non- canonical splicing junctions between exons and TEs (JETs). Some JETs are only detected in tumors, and not or rarely detected across large collections of healthy tissue samples.SUMMARY

[0005] The present disclosure identifies recurrent JETs that can generate novel protein isoforms that are stable products that contribute to the cellular functional proteome.

[0006] The present disclosure provides method for identifying a novel protein isoform, wherein the novel protein isoform comprises non-exonic amino acid sequence, comprising the steps of (a) aligning RNA sequences reads from a mammalian cell or tissue to a reference mammalian genome, (b) assembling transcripts from the genome-aligned RNA sequence reads of step (a), optionally using the combination of STAR and StringTie, (c) selecting variant transcripts from step (b) that (i) comprise one or more mutations compared to a reference mammalian genome, (ii) comprise exonic sequence, and (iii) comprise non-exonic sequence, optionally wherein the non-exonic sequence is a transposable element (TE), (d) optionally selecting variant transcripts from step (b) or step (c) that contain junctions between exonic sequence and non-exonic sequence and optionally are recurrently expressed in at least 1 % of subjects from a population of subjects suffering from a disease, optionally cancer, and (e) identifying the open reading frames (ORFs) of the variant transcripts from step (c) or (d) that are translated into protein, thereby identifying the novel protein isoform. Such a method for identifying the novel protein isoform may further comprise the step of selecting variant transcripts present in mRNA subpopulations of a mammalian cell or tissue identified as directly bound to ribosomes, optionally using RiboSeq sequencing, RiboseQC and / or ORFquant; and optionally comprising the step of selecting the ORFs from step (e) that are present in peptide fragments identified through mass spectrometry proteomics. In such methods, the non-exonic region of the reference mammalian genome can be a TE, optionally intronic TE, optionally a short interspersed nuclear element (SINE), or a long interspersed nuclear element (LINE), or a long terminal repeat (LTRs). In such methods, the exonic sequence can be from a gene of phylostratum 1.

[0007] The present disclosure provides novel protein isoforms, fragments thereof, and polynucleotides encoding such protein isoforms or fragments thereof; neoantigenic peptides of the novel protein isoforms; polynucleotides encoding such protein isoforms, fragments, or peptides, optionally linked to one or more heterologous regulatory control nucleotide sequences; vectors comprising such polynucleotides; oligonucleotide therapeutic agents targeting such polynucleotides; vaccine or immunogenic compositions comprising such protein isoforms or fragments thereof or neoantigenic peptides thereof, or such polynucleotides; an antibody, or an antigen-binding fragment thereof, a T cell receptor (TCR) in particular a non-HLA restricted TCR, or a chimeric antigen receptor (CAR) that specifically binds such protein isoforms or fragments thereof; methods of producing such antibodies, TCRs or CARs; polynucleotides encoding such antibodies, CARs or TCRs, optionally linked to one or more heterologous regulatory control nucleotide sequences; host cells, dendritic cells, antigen- presenting cells (APC) or immune cells that specifically bind to such protein isoforms or fragments thereof; and methods of using such products, including diagnostic or therapeutic methods.

[0008] The present disclosure provides an isolated protein isoform identified by the methods described herein, encoded by a variant transcript identified by the methods herein, or a fragment thereof at least 4, 5, 6, 7 or 8 amino acids in length. More particularly, the present disclosure provides a novel protein isoform comprising or consisting of any of the amino acid sequences of Table 1 A, or a fragment thereof at least 4, 5, 6, 7 or 8 amino acids in length. Some of these sequences, e.g. the sequences listed in Table IB, contain transmembrane domains and are expressed on the cell surface. The present disclosure also provides an isolated protein isoform, or neoantigenic peptides thereof, that comprise non-exonic amino acid sequence, optionally non-exonic amino acid sequence (a) that overlaps a junction of exonic amino acid sequence and non-exonic amino acid sequence, or (b) that comprises sequence encoded by a portion of a TE, or (c) that is encoded by a frameshifted (non-canonical) ORF downstream of the junction of exonic sequence and non-exonic sequence; or a fragment thereof at least 4, 5, 6, 7 or 8 amino acids in length. In some embodiments, the neoantigenic peptide is a fragment about 8 to about 16 amino acids in length of the protein isoform of any of claims 4-6, that comprises such non- exonic amino acid sequence. When the protein is a transmembrane protein, the neoantigenic peptide can be from the extracellular portion of the novel protein isoform.

[0009] In any of these embodiments, protein isoform can be encoded by a variant transcript that comprises non-exonic sequence, for example, a TE, optionally intronic TE, optionally a short interspersed nuclear element (SINE), or a long interspersed nuclear element (LINE), or a long terminal repeat (LTRs). In some embodiments, the protein isoform is encoded by variant transcript that comprises exonic sequence from a gene of phylostratum 1. In some embodiments, the non-exonic amino acid sequence is encoded by a portion of a reference mammalian genome that is not protein-coding sequence (e.g., encoded by intronic sequence, or encoded by an integrated transposable element). In some embodiments, the non-exonic amino acid sequence is encoded by nucleotide sequence comprising or overlapping the junction (or breakpoint) between exonic sequence and non-exonic sequence. In some embodiments, thenon-exonic amino acid sequence is from a non-canonical open reading frame (ORF) that results from a frameshift relative to the canonical open reading frame, where the frameshift is downstream of the junction between exonic sequence and non-exonic sequence.

[0010] Some of the novel protein isoforms, e.g., the sequences listed in Tables 3 and 4, are associated with cancer. Some of the novel protein isoforms, e.g. the sequences listed in Table 3, are correlated with patient survival after cancer. Typically, the novel protein isoform, fragment thereof, or neoantigenic peptide thereof, is expressed in more than 1%, notably more than 5 %, and typically more than 10% of the tumor samples. Typically the novel protein isoform, fragment thereof, or neoantigenic peptide thereof, is expressed at higher levels in tumor samples as compared to normal samples. Typically, the novel protein isoform, fragment thereof, or neoantigenic peptide thereof, is expressed in less than 20%, notably less than 10 %, less than 5 % or less than 1 % of the normal samples.

[0011] The present disclosure further provides a polynucleotide encoding any of the novel protein isoforms, fragments thereof, or neoantigenic peptides thereof described herein. In some embodiments, the polynucleotide sequence is a fragment at least 15, 20, 25 or 30 nucleotides in length that encodes non-exonic amino acid sequence. More particularly, the present disclosure provides polynucleotides that encode the amino acid sequence of any of the amino acid sequences of Table 1A, or a fragment thereof at least 15, 20, 25 or 30 nucleotides in length that encodes non-exonic amino acid sequence. In some embodiments, the polynucleotide sequence comprises any of the protein-coding portions of the nucleotide sequences of Table 1A or a fragment thereof. The polynucleotide sequence can be, e.g., DNA or RNA, or chemically modified versions thereof.

[0012] The present disclosure also provides an expression vector comprising any of the polynucleotides described herein, including fragments thereof at least 15, 20, 25 or 30 nucleotides in length, operably linked to one or more heterologous regulatory control nucleotide sequences. Examples of heterologous regulatory control nucleotide sequence include a promoter, a transcriptional transactivator, an enhancer, a translation optimizing sequence, and / or a polyadenylation signal. The present disclosure also provides a host cell comprising such expression vectors, and methods of using such host cells to produce the novel protein isoform, or fragment thereof.

[0013] The disclosure further provides a population of dendritic cells or antigen presenting cells, optionally autologous, that have been pulsed with any of the novel protein isoforms,fragments thereof, or neoantigenic peptides thereof described herein, or that have been transfected with a polynucleotide encoding any of the novel protein isoforms, fragments thereof, or neoantigenic peptides thereof described herein, or an expression vector comprising any of such polynucleotides.

[0014] The present disclosure further provides a vaccine or immunogenic composition capable of raising a specific immune cell response comprising (a) any of the novel protein isoforms, fragments thereof, or neoantigenic peptides thereof described herein, optionally with a physiologically acceptable buffer, carrier, or excipient, and / or optionally with an adjuvant or immunostimulant; (b) a polynucleotide encoding any of the novel protein isoforms, fragments thereof, or neoantigenic peptides thereof as described herein, or an expression vector comprising such polynucleotide, or (c) a population of antigen presenting cells as described herein.

[0015] The present disclosure also provides an oligonucleotide therapeutic agent that targets a polynucleotide encoding any of the novel protein isoforms, or fragments thereof described herein, and preferably does not target the canonical gene transcript. More particularly, the oligonucleotide therapeutic agent may target a polynucleotide encoding any of the amino acid sequences of Table 1A, and does not target the canonical gene transcript. An oligonucleotide therapeutic agent can target any polynucleotide comprising a nucleotide sequence encoding any of the amino acid sequences of Table 1A, optionally a nucleotide sequence of Table 1A, and which does not target the canonical gene transcript. The oligonucleotide therapeutic agent is preferably at least 10, 15, 20, 25 or 30 nucleotides in length. The oligonucleotide therapeutic agent can comprise (i) a nucleotide sequence that binds to or is complementary to a fragment of any of the nucleotide sequences of Table 1A, or wherein said fragment optionally comprises a junction of exonic sequence and non-exonic sequence or about 40 to about 200 nucleotides of adjacent sequence, or (ii) a nucleotide sequence that is complementary to a region comprising a putative splicing site of any of the nucleotide sequences of Table 1A.

[0016] For example, the oligonucleotide therapeutic agent comprises a nucleotide sequence that is at least 10, 12, 15, 20, 25 or 30 nucleotides in length that binds to or is complementary to a polynucleotide comprising a protein-coding portion encoding any of the amino acid sequences of Table 1 A, or a fragment thereof that comprises a junction of exonic sequence and non-exonic sequence. In some embodiments, the oligonucleotide therapeutic agent comprises a nucleotide sequence that is at least 10, 12, 15, 20, 25 or 30 nucleotides in length that binds to or is complementary to any of the nucleotide sequences of Table 1A or a fragment thereof thatcomprises a junction of exonic sequence and non-exonic sequence. In some embodiments, the oligonucleotide therapeutic agent comprises a nucleotide sequence that is at least 10, 12, 15, 20, 25 or 30 nucleotides in length that binds to or is complementary to a fragment of the nonprotein coding portion of any of the nucleotide sequences of Table 1A, or adjacent sequence, for example, upstream or downstream sequence comprising the promoter or 5’ or 3’ regulatory control sequence. Such oligonucleotide therapeutic agent may increase or decrease expression or level of the protein isoform, or of the canonical gene. In some embodiments, the oligonucleotide therapeutic agent comprises a nucleotide sequence that is at least 10, 12, 15, 20, 25 or 30 nucleotides in length that binds to or is complementary to a region comprising a putative splicing site of any of the nucleotide sequences of Table 1A. Such oligonucleotide therapeutic agent may induce alternative splicing that restores a portion of the canonical exons. In some embodiments, the oligonucleotide therapeutic agent comprises a nucleotide sequence that is at least 10, 12, 15, 20, 25 or 30 nucleotides in length that binds to or is complementary to a fragment of the canonical gene corresponding to any of the nucleotide sequences of Table 1A or adjacent sequence. Such oligonucleotide therapeutic agent may increase or decrease expression or level of the canonical gene. Suitable oligonucleotide therapeutic agents are known in the art and include antisense RNA, RNAi, shRNA, microRNA, saRNA, aptamers, ribozymes, SSOs. In some embodiments, the oligonucleotide therapeutic agent targets any of the nucleotide sequences of PTEN JET-ORF, optionally for use in inducing an inflammatory response, or increasing IL-6 expression. In some embodiments, the oligonucleotide therapeutic agent targets any of the nucleotide sequences of WWOX JET-ORF, optionally for use in the modulation of ATF6, KLF5, TWIST1 and TP73 transcriptional responses.

[0017] The present disclosure further encompasses an antibody, or antigen-binding fragment thereof, or antigen binding domain thereof, that binds any of the novel protein isoforms, fragments thereof, or neoantigenic peptides thereof described herein, preferably that comprise a non-exonic amino acid sequence as described herein. More particularly, the antibody specifically binds any of the amino acid sequences of Table 1A or a fragment thereof that comprises a non-exonic amino acid sequence (e.g., comprises the junction of exonic and non- exonic amino acid sequence, or comprises non-canonical amino acid sequence due to a frameshifted ORF or alternative splicing event). In some embodiments, the antibody binds a protein isoform that comprises a transmembrane domain, e.g. is expressed on the cell surface. The antibody, or antigen-binding domain or fragment thereof, can bind with a Kd binding affinity of about 10'6M or less (lower numbers indicating higher binding affinity), or about 10"7M or less, or about 10'8M or less, or about 10'9M or less, or about IO'10M or less, or about 10'11M or less, or about 10'12M or less. In some embodiments, the antibodies, TCRs or CARs that specifically bind the novel protein isoforms, fragments thereof, or neoantigenic peptides thereof described herein may bind a sequence of at least 4, at least 5, at least 6, or at least 7 amino acids. Typically, the antigen binding fragment binds a non-exonic amino acid sequence.

[0018] In certain embodiments, the antigen binding domain comprises one or more, typically one or two immunoglobulin region(s). The antigen binding domain can comprise a heavy chain variable region (VH) of an antibody and / or a light chain variable region (VL) of an antibody. The antigen-binding fragment can comprise at least one, two or preferably three CDRs of the antibody.

[0019] The present disclosure also encompasses an antibody comprising an antigen binding domain as herein defined wherein the antibody is selected from a full IgG, an scFv, a BiTE, or a multispecific antibody. The antibody can be of human, murine or camelid origin.

[0020] The present disclosure also encompasses a chimeric antigen receptor (CAR) or a non- HLA restricted recombinant T cell receptor (TCR) comprising an antigen-binding fragment as described herein.

[0021] Typically non-HLA restricted recombinant TCR of the present disclosure comprises an extracellular antigen-binding domain which is capable of dimerizing with a second extracellular antigen-binding domain. In some embodiments, the second extracellular antigen-binding domain binds a tumor antigen, preferably wherein the tumor antigen is selected from pHER95, CD 19, MUC16, MUC1, CAIX, CEA, CD8, CD7, CD 10, CD20, CD22, CD30, CD70, CLL1, CD33, CD34, CD38, CD41, CD44, CD49f, CD56, CD74, CD133, CD138, EGP-2, EGP-40, EpCAM, Erb-B2, Erb-B3, Erb-B4, FBP, Fetal acetylcholine receptor, folate receptor-a, GD2, GD3, HER-2, hTERT, IL-13R-a2, K-light chain, KDR, LeY, LI cell adhesion molecule, MAGE-A1, Mesothelin, MAGEA3, p53, MARTI, GP100, Proteinase3 (PR1), Tyrosinase, Survivin, hTERT, EphA2, NKG2D ligands, NY-ESO-1, oncofetal antigen (h5T4), PSCA, PSMA, ROR1, TAG-72, VEGF-R2, WT-1, BCMA, CD123, CD44V6, NKCS1, EGF1R, EGFR-VIII, CD99, CD70, ADGRE2, CCR1, LILRB2, LILRB4, PRAME, and ERBB.

[0022] The present disclosure further encompasses a CAR comprising:

[0023] a) an extracellular region comprising the antigen-binding domain or antigen-binding fragment described herein,

[0024] b) a transmembrane domain,

[0025] c) optionally one or more costimulatory domains

[0026] d) an intracellular signaling domain comprising a modified CD3zeta intracellular signaling domain in which ITAM2 and ITAM3 have been inactivated,

[0027] In some embodiments, the CAR as herein disclosed comprises a transmembrane domain selected from CD28, CD8 or CD3-zeta.

[0028] In some embodiments, the CAR as herein disclosed comprises one or more costimulatory domains which can be selected from the group consisting of 4-1BB, CD28, ICOS, 0X40 and DAP 10.

[0029] In some embodiments, the CAR as herein disclosed comprises an intracellular signaling domain comprising the intracellular signaling domain of a CD3-zeta polypeptide, or a fragment thereof, optionally a CD3-zeta polypeptide wherein immunoreceptor tyrosine-based activation motif 2 (ITAM2) and immunoreceptor tyrosine-based activation motif 3 (ITAM3) are inactivated.

[0030] The present disclosure also encompasses: an antibody, or an antigen-binding fragment thereof, a T cell receptor (TCR), optionally that has been selected for its binding affinity to any of the novel protein isoforms, or fragments thereof or neoantigenic peptides thereof described herein, e.g. of any of the amino acid sequences of Table 1 A, or a fragment thereof, optionally of a length at least 4, 5, 6 7, or 8 amino acids, or a non-HLA restricted recombinant T cell receptor (TCR), or a chimeric antigen receptor (CAR) comprising an antigen-binding fragment or antigen-binding domain of such antibody or TCR; a composition comprising such antibody, antigen-binding fragment thereof, TCR or CAR, optionally in combination with a physiologically or pharmacologically acceptable buffer, preferably sterile; or a polynucleotide encoding an antibody, a CAR or a TCR as herein defined, typically operatively linked to a heterologous regulatory control nucleotide sequence, and a vector encoding such polynucleotide.

[0031] The present disclosure further provides an immune cell, or a population of immune cells that targets any of the novel protein isoforms, or fragments thereof or neoantigenic peptides thereof described herein, e.g. of any of the amino acid sequences of Table 1A, e.g. of a lengthat least 4, 5, 6 7, or 8 amino acids, wherein the population of immune cells preferably targets a plurality of different novel protein isoforms, or fragments thereof or neoantigenic peptides described herein, e.g. of any of the amino acid sequences of Table 1A. Typically the immune cell or population of immune cells comprises any of the antibodies, antigen-binding domains or antigen-binding fragments, or TCRs or CARs described herein, or comprises any of the polynucleotides encoding an antibody, CAR or TCR disclosed herein, optionally wherein the polynucleotide is operatively linked to a heterologous regulatory control nucleotide sequence. Also provided is a composition comprising such immune cells or population of immune cells, optionally in combination with a physiologically or pharmacologically acceptable buffer, carrier, excipient, immunostimulant and / or adjuvant, preferably sterile.

[0032] In some embodiments, the antibody or antigen-binding fragment thereof, TCR or CAR binds a novel protein isoform, or fragment thereof or neoantigenic peptide thereof described herein that is expressed on the surface of a cell, with a Kd affinity of about 10'6M or less or about 10'7M or less. In some embodiments, the antibody or antigen-binding fragment thereof, or TCR or CAR is a multispecific, for example, it further targets at least one immune cell antigen. In some embodiments, the TCR is a soluble fragment of a TCR fused to an antibody fragment that targets at least one immune cell antigen. The immune cell antigen can be from a T cell, a NK cell or a dendritic cell, and optionally the immune cell antigen is CD3, CD16, CD30 or a TCR.

[0033] The present disclosure further provides a method of producing an antibody, a TCR, e.g., non-HLA restricted TCR, or a CAR as herein defined comprising the step of selecting an antibody, TCR or a CAR that binds to any of the novel protein isoforms, or fragments thereof or neoantigenic peptides thereof described herein, typically a protein isoform of any one of the amino acid sequences of Table 1A, and that preferably specifically binds the novel protein isoform, or fragment thereof or neoantigenic peptide thereof, e.g., with a Kd affinity constant of about 10'6M or less, or about 10'7M or less. In some embodiments, the antibody, TCR, or CAR is selected after contacting a population of candidate antibodies, TCRs or CARs with said protein isoform or fragment or neoantigenic peptide. In some embodiments, the antibody, TCR or CAR is selected from a population produced by immunizing an animal with said protein isoform or fragment or neoantigenic peptide as disclosed herein.

[0034] The present disclosure also provides an ex vivo method for producing a T cell comprising a TCR that specifically binds any of the novel protein isoforms, or fragments thereof or neoantigenic peptides thereof, comprising: (a) stimulating T cells from a blood sampleobtained from a patient with a composition comprising said protein isoform, fragment or peptide, or polynucleotide encoding said protein isoform, fragment or peptide, to prime, activate, and expand T-cells, optionally CD4+ T-cells, CD8+ T cells, or effector or central memory T cells, (b) optionally prior to step (a), the method comprises obtaining a sample of cells or tissue from the patient, optionally by leukapheresis or tumor biopsy, and confirming expression of said protein isoform or fragment thereof in said sample, (c) optionally subsequent to step (a), the method comprises confirming the specificity and functionality, optionally cytokine production activity and cytolytic activity, of the induced T-cells. Such T-cells are useful in cell therapy, e.g., cell therapy for treating cancer.

[0035] The present disclosure further encompasses a polynucleotide encoding an antibody, a CAR or a non-HLA restricted TCR as herein defined, optionally linked to a heterologous regulatory control nucleotide sequence and vectors comprising thereof. The present disclosure also encompasses an immune cell comprising an antigen-binding fragment of an antibody, a CAR or TCR or a polynucleotide encoding such an antigen-binding fragment of an antibody, a CAR or TCR, in particular a non-HLA restricted TCR as defined herein. Said immune cell can be allogenic or autologous. It is typically selected from T cells, Natural Killer T cells, CD4+ T cells, CD8+ T cells, CD4+ / CD8+ T cells, TILs / tumor derived CD8 T cells, central memory CD8+ T cells, Treg, Mucosal-Associated Invariant T cells (MAIT), alpha / beta or gamma / delta T cells, human embryonic stem cells, pluripotent stem cells, and / or myeloid or lymphoid lineage cells. In certain embodiment, the immune cell is defective for Suv39hl, in particular in said immune cell the Suv39hl gene is disrupted by deletion of the entire gene, exon, or region, replacement with an exogenous sequence, and / or mutation by frameshift or missense mutation within the gene suv39hl gene. The Suv39hl gene is typically the human Suv39hl gene encoding the human Suv39hl protein referenced 043463 in UniProt. Methods of preparing such immune cells are also contemplated, for example, by delivering a nucleic acid or vector encoding any of the antibody, TCR, or CAR described herein to the cell, in vivo or ex vivo.

[0036] The present disclosure also encompasses diagnostic and therapeutic uses of such pharmaceutical compositions, novel protein isoforms, or fragments thereof or neoantigenic peptides thereof; polynucleotides or vectors encoding such proteins, fragments or peptides thereof; oligonucleotide therapeutic agents; vaccine or immunogenic compositions; host cells or dendritic cells or APCs; antibody, antigen-binding fragment thereof, TCR or CAR; or immune cells described herein. Such therapeutic uses include in particular for inhibiting cancer cell proliferation or for cancer treatment in a subject in need thereof. Such therapeutic usesinclude for use in treating a disease associated with the protein isoform sequence in Table 1 A, 3, or 4.

[0037] In particular, uses of the oligonucleotide therapeutic agents include for use in modulating expression of and / or controlling function of any of the protein isoforms described herein, optionally for inhibiting expression of the protein isoform. In such uses, the oligonucleotide therapeutic agents are preferably specific for the protein isoforms described herein, e.g. bind to or are complementary to the non-exonic sequence or the junction of the exonic and non-exonic sequence. Alternatively, the oligonucleotide therapeutic agents are for use in modulating expression of and / or controlling function of the canonical protein associated with the protein isoform sequence. Any of such uses include uses for treating a disease associated with the protein isoform sequence in Table 1 A, 3, or 4.

[0038] Typically, the composition further comprises a pharmaceutical excipient. Treatment as used herein includes both prophylactic and therapeutic treatment.

[0039] The present disclosure also encompasses the use in cell therapy in a subject in need thereof, for example, cell therapy of cancer, of pharmaceutical compositions, novel protein isoforms, or fragments thereof or neoantigenic peptides thereof; polynucleotides or vectors encoding such proteins, fragments or peptides thereof; vaccine or immunogenic compositions; host cells or dendritic cells or APCs; antibody, antigen-binding fragment thereof, TCR or CAR; or immune cells described herein. Typically, the composition further comprises a pharmaceutical excipient, preferably sterile.

[0040] The disclosure also provides pharmaceutical compositions comprising an effective amount of any of the novel protein isoforms, or fragments thereof or neoantigenic peptides thereof; polynucleotides or vectors encoding such proteins, fragments or peptides thereof; oligonucleotide therapeutic agents; vaccine or immunogenic compositions; host cells or dendritic cells or APCs; antibody, antigen-binding fragment thereof, TCR or CAR; or an immune cell as described herein; together with a pharmaceutically acceptable excipient, optionally with a sterile pharmaceutically acceptable excipient(s), carrier, and / or buffer; as well as methods of using them.

[0041] In any of the embodiments described herein, pharmaceutical compositions can be administered in combination with at least one further therapeutic agent. Such further therapeutic agent, particularly for cancer, can typically be a chemotherapeutic agent, or an immunotherapeutic agent, optionally a checkpoint inhibitor.

[0042] For example, according to the present disclosure, any of the pharmaceutical compositions can be administered in combination with an anti- immunosuppressive / immunostimulatory agent. For example, the subject is further administered with one or more checkpoint inhibitors typically selected from PD-1 inhibitors, PD-L1 inhibitors, Lag-3 inhibitors, Tim-3 inhibitors, TIGIT inhibitors, BTLA inhibitors, V-domain Ig suppressor of T-cell activation (VISTA) inhibitors and CTLA-4 inhibitors, or IDO inhibitors.

[0043] The present disclosure also provides diagnostic methods, for example, a method for determining the prognosis of a cancer or for detecting cancer cells characterized by increased expression of a protein isoform of Table 3 or 4. In some embodiments, the method comprises the step of detecting or quantifying a protein isoform described herein, e.g. a protein isoform of any of the amino acid sequences of Table 1 A, in a biological sample from a patient, e.g. by contacting the biological sample with an antibody, including an antigen-binding fragment thereof, that specifically binds the protein isoform, and detecting binding of the antibody with the protein isoform. In some embodiments, the method comprises the step of detecting or quantifying a nucleic acid encoding any of the amino acid sequences of Table 1 A in a biological sample isolated from a patient, optionally wherein the detecting comprises: (a) contacting said nucleic acid from the biological sample with a probe which specifically hybridizes to any of the nucleotide sequences of Table 1A, preferably a probe that hybridizes to the junction of exonic sequence and non-exonic sequence; (b) optionally detecting a complex formed between the probe and the nucleic acid from the biological sample, and comparing the amount of the complex in the biological sample to the amount of the complex in a reference sample comprising corresponding normal tissue; (c) wherein detecting said complex in the biological sample is associated with a better or worse prognosis of cancer, (d) or wherein detecting an increased amount of said complex in the biological sample as compared to the corresponding normal tissue indicates that the biological sample comprises cancer cells. In particular, the protein isoforms of Table 4 are associated with cancer, and the protein isoforms of Table 3 are associated positively or negatively with cancer survival.

[0044] Various embodiments of the foregoing methods and products and compositions are described in detailed below. Except for alternatives clearly mentioned, combinations of such embodiments are encompassed by the present application.FIGURE LEGENDFigure 1: Schematic representation of the method for identifying novel and stably expressed splice variants containing TE / exon junctions.Figure 2: Spliced RNA-seq reads mapping to both a protein-coding exon and a TE of two examples of identified JETs.Figure 3: The number of JETs per sample.Figure 4: JET patient recurrence distribution in TCGA cancer types.Figure 5: JET patient recurrence distribution in LU AD and LGG.Figure 6: . UMAP visualization based on JET expression.Figure 7: Recurrent JETs in tumor-adjacent normal tissues from TCGA.Figure 8: Correlation of JET recurrence in TCGA tumors and in normal tissues.Figure 9: JET position within ORF.Figure 10: JET and reference ORF length in amino acids.DETAILED DISCLOSURE

[0045] The present disclosure identifies recurrent JETs that can generate stable products that contribute to the cellular functional proteome. RiboSeq and mass spectrometry (MS)-based proteomics were employed to identify recurrent and translated TE exonization events. A population of low abundance protein-encoding variant transcripts was identified that are translated as efficiently as canonical isoforms. Data herein shows that some of these protein isoforms are stable in cells, localize to defined subcellular compartments, and present modified functions compared to canonical isoforms. These identified variant transcripts comprise exonic sequence and non-exonic sequence. In some embodiments, the non-exonic sequence is derived from a transposable element (TE). The non-exonic sequence can be 5’ or 3’ of the exonic sequence.

[0046] Non-exonic amino acid sequence can be encoded by a portion of a reference mammalian genome that is not protein-coding sequence (e.g., encoded by intronic sequence, or encoded by an integrated transposable element). Non-exonic amino acid sequence can be encoded by nucleotide sequence comprising or overlapping the junction (or breakpoint) between exonic sequence and non-exonic sequence. Non-exonic amino acid sequence can alternatively result from a frameshift relative to the canonical open reading frame (ORF), where the frameshift is downstream of the junction between exonic sequence and non-exonic sequence and results inexpression of non-canonical exons. Non-exonic amino acid sequence can also result from introduction of a novel start codon that results in expression of upstream non-protein-coding sequence, thus introducing additional protein sequence or additional exons compared to the canonical protein sequence. Non-exonic amino acid sequence can also result from alternative splicing events.

[0047] In one aspect of the disclosure, provided herein is a novel method for selecting relatively low abundance isoforms of proteins or polypeptides and neoantigenic peptides that are encoded by a variant transcript sequence comprising a part of a non-exonic sequence, e.g. TE sequence, and a part of an exonic sequence. The novel protein isoform identified by these methods includes exonic amino acid sequence and non-exonic amino acid sequence.

[0048] Some of these retain the function of the canonical isoform, while others have altered function that is linked to disease. Some of these are transmembrane proteins that represent diagnostic or therapeutic targets for antibodies and other binding partners. The present disclosure also allows selecting peptides having shared isoforms among a population of patients. Such shared isoforms are of high therapeutic interest since they can be used in diagnosis or therapy for a large population of patients.

[0049] In another aspect of the disclosure, provided herein are neoantigenic protein, polypeptide and peptide sequences, that preferably comprise non-exonic amino acid sequence. In some embodiments, the protein, polypeptide or peptide sequences overlap or comprise the junction between the exonic sequence and non-exonic sequence. The disclosure also provides polynucleotides and expression vectors that encode these protein, polypeptide and peptide sequences. The neoantigenic peptides identified by the method according to the present disclosure are typically highly immunogenic, particularly when they are derived from variant transcripts absent from normal cells. Also provided herein are vaccines comprising the neoantigenic protein, polypeptide and peptide sequences, or the encoding nucleic acids, optionally as part of expression vectors. Further provided herein are antigen-presenting cells (APC) that present the peptide sequences disclosed herein, preferably bound to at least one Major Histocompatibility Complex (MHC) molecule.

[0050] In a further aspect of the disclosure, provided herein are oligonucleotide therapeutic agents such as antisense RNA, RNAi, shRNA, microRNA, saRNA, aptamers, ribozymes, or SSOs, that bind to the variant transcripts. In some embodiments, the oligonucleotidetherapeutic agents specifically or preferentially bind to the variant transcripts compared to the canonical transcripts.

[0051] In yet another aspect of the disclosure, provided herein are antibodies and antigenbinding fragments thereof, recombinant T-cell receptors (TCRs) and antigen-binding fragments thereof, and other natural or synthetic binding partners, that bind to the neoantigenic protein, polypeptide and peptide sequences of the disclosure. Also provided herein are methods of producing or screening for the antibodies, recombinant TCR and binding partners. Further provided are immune cells, including antigen-presenting cells (APC), T-cells, NK cells, B cells, and lymphoid progenitor cells, myeloid progenitor cells, and stem cells, that comprise such antibodies, antigen-binding fragments thereof, recombinant T-cell receptors (TCRs), antigenbinding fragments thereof, or other natural or synthetic binding partners.

[0052] Examples of genes identified herein for which novel protein isoforms have been identified include:

[0053] H2AFY gene encodes for a variant of H2A histone, and it has been found to be dysregulated in various tumors. The H2AFY isoform identified herein is thus a potential biomarker for unfavorable prognosis and correlates with immune infiltration in HCC. See Huang et al., “Exploring the Prognostic Value, Immune Implication and Biological Function of H2AFY Gene in Hepatocellular Carcinoma,” Frontiers in Immunol., vol. 12 (24 Nov 2021).

[0054] The ALDH3A2 isoform identified herein can be useful as a potential reference value for the relief and immunotherapy of gastric adenocarcinoma, and also as an independent predictive marker for prognosis of gastric adenocarcinoma. See Yin et al., “Identification of ALDH3A2 as a novel prognostic biomarker in gastric adenocarcinoma using integrated bioinformatics analysis,” BMC Cancer volume 20, Article number: 1062 (2020).

[0055] The IL15RA isoform identified herein can be correlated with poor clinical outcome in cancer. Expression of IL-15 and / or IL-15 receptors can be detected in several leukemia-derived and solid tumor-derived cell lines, where they display protumorigenic properties. Increased levels of sIL-15Ra in the serum of patients with head and neck cancer and high intratumoral IL-15 concentration, are highly correlated with poor clinical outcome. See Fiore et al., “Interleukin- 15 and cancer: some solved and many unsolved questions,” Journal for ImmunoTherapy of Cancer; 8 :e001428 (2020).

[0056] The PTEN isoform herein can also be a potential biomarker for cancer and unfavorable prognosis in cancer. PTEN (Phosphatase and TENsin homolog) tumor suppressor gene (basedon lipid phosphatase activity) is often mutated in malignant tumor with high incidence in the general population. See Eng, “ PTEN: one gene, many syndromes,” Hum Mutat, 22: 183-981 (2003); Trotman et al, “Ubiquitination regulates PTEN nuclear import and tumor suppression,” Cell 128: 141-56 (2007); Shen et al., “Essential role for nuclear PTEN in maintaining chromosomal integrity,” Cell, 128 : 157-70 (2007).

[0057] The WWOX isoform herein can also be a potential biomarker for cancer and unfavorable prognosis in cancer. WWOX tumor suppressor activity has been proposed on a basis of numerous genomic alterations reported in chromosome 16q23.3-24.1 locus. The WWOX gene is localized on the long arm of chromosome 16, locus 23.3-24.1 in humans. It spans the one of the most frequently expressed common human chromosomal fragile site, FRA16D, which is involved in numerous cancers, such as breast, ovarian, prostate, lung, esophageal, gastric, pancreatic, and hepatic cancer. As the FRA16D site is highly susceptible to DNA damage, a high rate of translocations, deletions, increased sister chromatid exchanges, and other alterations are observed within it. Some alternative splicing variants are produced by deletions of exons 5-8 and 6-8 in WWOX mRNA. See Baryla et al., Exp Biol Med (Maywood), “Alteration of WWOX in human cancer, a clinical view,” 240(3): 305-314 (2015).Definitions

[0058] According to the present disclosure, the term “exonic sequence” refers to nucleotide sequence from protein-coding sequence of a reference mammalian genome (also referred to herein as sequence from a canonical exon). The term “non-exonic sequence” refers to nucleotide sequence from a portion of a reference mammalian genome that is not protein-coding sequence or that is TE sequence. Reference mammalian genome can be human genome.

[0059] The term “exonic amino acid sequence” refers to amino acid sequence encoded by protein-coding sequence of a reference mammalian genome (i.e., amino acid sequence from a canonical exon). The term “non-exonic amino acid sequence” refers to amino acid sequence that is not part of a canonical exon of the reference mammalian genome.

[0060] The term “disease” refers to any pathological state, including cancer diseases, in particular those forms of cancer diseases described herein.

[0061] The term “normal” refers to the healthy state or the conditions in a healthy subject or tissue, i.e., non-pathological conditions, wherein “healthy” preferably means non-cancerous.

[0062] Cancer (medical term: malignant neoplasm) is a class of diseases in which a group of cells display uncontrolled growth (division beyond the normal limits), invasion (intrusion on and destruction of adjacent tissues), and sometimes metastasis (spread to other locations in the body via lymph or blood). These three malignant properties of cancers differentiate them from benign tumors, which are self-limited, and do not invade or metastasize. Most cancers form a tumor but some, like leukemia, do not.

[0063] Malignant tumor is essentially synonymous with cancer. Malignancy, malignant neoplasm, and malignant tumor are essentially synonymous with cancer.

[0064] As used herein, the term “tumor” or “tumor disease” refers to an abnormal growth of cells (called neoplastic cells, tumorigenous cells or tumor cells) preferably forming a swelling or lesion. By “tumor cell” is meant an abnormal cell that grows by a rapid, uncontrolled cellular proliferation and continues to grow after the stimuli that initiated the new growth cease. Tumors show partial or complete lack of structural organization and functional coordination with the normal tissue, and usually form a distinct mass of tissue, which may be either benign, pre- malignant or malignant.

[0065] A benign tumor is a tumor that lacks all three of the malignant properties of a cancer. Thus, by definition, a benign tumor does not grow in an unlimited, aggressive manner, does not invade surrounding tissues, and does not spread to non-adjacent tissues (metastasize).

[0066] Neoplasm is an abnormal mass of tissue as a result of neoplasia. Neoplasia (new growth in Greek) is the abnormal proliferation of cells. The growth of the cells exceeds, and is uncoordinated with that of the normal tissues around it. The growth persists in the same excessive manner even after cessation of the stimuli. It usually causes a lump or tumor. Neoplasms may be benign, pre-malignant or malignant.

[0067] “Growth of a tumor” or “tumor growth” according to the present disclosure relates to the tendency of a tumor to increase its size and / or to the tendency of tumor cells to proliferate.

[0068] For purposes of the present disclosure, the terms “cancer” and “cancer disease” are used interchangeably with the terms “tumor” and “tumor disease”.

[0069] Cancers are classified by the type of cell that resembles the tumor and, therefore, the tissue presumed to be the origin of the tumor. These are the histology and the location, respectively.

[0070] According to the present application, cancer may affect any one of the following tissues or organs: breast; liver; kidney; heart, mediastinum, pleura; floor of mouth; lip; salivary glands; tongue; gums; oral cavity; palate; tonsil; larynx; trachea; bronchus, lung; pharynx, hypopharynx, oropharynx, nasopharynx; esophagus; digestive organs such as stomach, intrahepatic bile ducts, biliary tract, pancreas, small intestine, colon; rectum; urinary organs such as bladder, gallbladder, ureter; rectosigmoid junction; anus, anal canal; skin; bone; joints, articular cartilage of limbs; eye and adnexa; brain; peripheral nerves, autonomic nervous system; spinal cord, cranial nerves, meninges; and various parts of the central nervous system; connective, subcutaneous and other soft tissues; retroperitoneum, peritoneum; adrenal gland; thyroid gland; endocrine glands and related structures; female genital organs such as ovary, uterus, cervix uteri; corpus uteri, vagina, vulva; male genital organs such as penis, testis and prostate gland; hematopoietic and reticuloendothelial systems; blood; lymph nodes; thymus.

[0071] The term “cancer” according to the disclosure therefore comprises leukemias, seminomas, melanomas, teratomas, lymphomas, neuroblastomas, gliomas, rectal cancer, endometrial cancer, kidney cancer, adrenal cancer, thyroid cancer, blood cancer, skin cancer, cancer of the brain, cervical cancer, intestinal cancer, liver cancer, colon cancer, stomach cancer, intestine cancer, head and neck cancer, gastrointestinal cancer, lymph node cancer, esophagus cancer, colorectal cancer, pancreas cancer, ear, nose and throat (ENT) cancer, breast cancer, prostate cancer, cancer of the uterus, ovarian cancer and lung cancer and the metastases thereof. Examples thereof are lung carcinomas, mamma carcinomas, prostate carcinomas, colon carcinomas, renal cell carcinomas, cervical carcinomas, or metastases of the cancer types or tumors described above. The term cancer according to the present disclosure also comprises cancer metastases and relapse of cancer. Lung cancer includes small cell lung carcinoma (SCLC) and non-small cell lung carcinoma (NSCLC), squamous cell lung carcinoma, lung adenocarcinoma (LU AD), and large cell lung carcinoma.

[0072] By “metastasis” is meant the spread of cancer cells from its original site to another part of the body. The formation of metastasis is a very complex process and depends on detachment of malignant cells from the primary tumor, invasion of the extracellular matrix, penetration of the endothelial basement membranes to enter the body cavity and vessels, and then, after being transported by the blood, infiltration of target organs. Finally, the growth of a new tumor, i.e. a secondary tumor or metastatic tumor, at the target site depends on angiogenesis. Tumor metastasis often occurs even after the removal of the primary tumor because tumor cells or components may remain and develop metastatic potential. In one embodiment, the term“metastasis” according to the present disclosure relates to “distant metastasis” which relates to a metastasis which is remote from the primary tumor and the regional lymph node system.

[0073] The cells of a secondary or metastatic tumor are like those in the original tumor. This means, for example, that, if ovarian cancer metastasizes to the liver, the secondary tumor is made up of abnormal ovarian cells, not of abnormal liver cells. The tumor in the liver is then called metastatic ovarian cancer, not liver cancer.

[0074] A relapse or recurrence occurs when a person is affected again by a condition that affected them in the past. For example, if a patient has suffered from a tumor disease, has received a successful treatment of said disease and again develops said disease said newly developed disease may be considered as relapse or recurrence. However, according to the present disclosure, a relapse or recurrence of a tumor disease may but does not necessarily occur at the site of the original tumor disease. Thus, for example, if a patient has suffered from ovarian tumor and has received a successful treatment a relapse or recurrence may be the occurrence of an ovarian tumor or the occurrence of a tumor at a site different to ovary. A relapse or recurrence of a tumor also includes situations wherein a tumor occurs at a site different to the site of the original tumor as well as at the site of the original tumor. Preferably, the original tumor for which the patient has received a treatment is a primary tumor and the tumor at a site different to the site of the original tumor is a secondary or metastatic tumor.

[0075] By “treat” is meant to administer a compound or composition as described herein to a subject in order to prevent or eliminate a disease, including reducing the size of a tumor or the number of tumors in a subject; arrest or slow a disease in a subject; inhibit or slow the development of a new disease in a subject; decrease the frequency or severity of symptoms and / or recurrences in a subject who currently has or who previously has had a disease; and / or prolong, i.e. increase the lifespan of the subject. In particular, the term “treatment of a disease” includes curing, shortening the duration, ameliorating, preventing, slowing down or inhibiting progression or worsening, or preventing or delaying the onset of a disease or the symptoms thereof.

[0076] By “being at risk” is meant a subject, i.e. a patient, that is identified as having a higher than normal chance of developing a disease, in particular cancer, compared to the general population. In addition, a subject who has had, or who currently has, a disease, in particular cancer, is a subject who has an increased risk for developing a disease, as such a subject maycontinue to develop a disease. Subjects who currently have, or who have had, a cancer also have an increased risk for cancer metastases.

[0077] The therapeutically active agents, vaccines and compositions described herein may be administered via any conventional route, including by injection or infusion.

[0078] The agents described herein are administered in effective amounts. An “effective amount” refers to the amount which achieves a desired reaction or a desired effect alone or together with further doses or together with further therapeutic agents. In the case of treatment of a particular disease or of a particular condition, the desired reaction preferably relates to inhibition of the course of the disease. This comprises slowing down the progress of the disease and, in particular, interrupting or reversing the progress of the disease. The desired reaction in a treatment of a disease or of a condition may also be delay of the onset or a prevention of the onset of said disease or said condition.

[0079] An effective amount of an agent described herein will depend on the condition to be treated, the severity of the disease, the individual parameters of the patient, including age, physiological condition, size and weight, the duration of treatment, the type of an accompanying therapy (if present), the specific route of administration and similar factors. Accordingly, the doses administered of the agents described herein may depend on various of such parameters. In the case that a reaction in a patient is insufficient with an initial dose, higher doses (or effectively higher doses achieved by a different, more localized route of administration) may be used.

[0080] The pharmaceutical compositions as herein described are preferably sterile and contain an effective amount of the therapeutically active substance to generate the desired reaction or the desired effect.

[0081] The pharmaceutical compositions as herein described are generally administered in pharmaceutically compatible amounts and in pharmaceutically compatible preparation. The term “pharmaceutically compatible” refers to a nontoxic material which does not interact with the action of the active component of the pharmaceutical composition. Preparations of this kind may usually contain salts, buffer substances, preservatives, carriers, supplementing immunityenhancing substances such as adjuvants, e.g. CpG oligonucleotides, cytokines, chemokines, saponin, GM-CSF and / or RNA and, where appropriate, other therapeutically active compounds. When used in medicine, the salts should be pharmaceutically compatible.

[0082] As used herein, the term “nucleic acid molecules” include any nucleic acid molecule that encodes a polypeptide of interest or a fragment thereof. Such nucleic acid molecules need not be 100% homologous or identical with an endogenous nucleic acid sequence but may exhibit substantial identity. Polynucleotides having “substantial identity” or “substantial homology” to an endogenous sequence are typically capable of hybridizing with at least one strand of a double-stranded nucleic acid molecule. By “hybridize” is meant a pair to form a double-stranded molecule between complementary polynucleotide sequences (e.g., a gene described herein), or portions thereof, under various conditions of stringency. (See, e.g., Wahl, G. M. and S. L. Berger (1987) Methods Enzymol. 152:399; Kimmel, A. R. (1987) Methods Enzymol. 152:507). For example, stringent salt concentration will ordinarily be less than about 750 mM NaCl and 75 mM trisodium citrate, e.g., less than about 500 mM NaCl and 50 mM trisodium citrate, or less than about 250 mM NaCl and 25 mM trisodium citrate. Low stringency hybridization can be obtained in the absence of organic solvent, e.g., formamide, while high stringency hybridization can be obtained in the presence of at least about 35% formamide, e.g., at least about 50% formamide. Stringent temperature conditions will ordinarily include temperatures of at least about 30° C, at least about 37° C, or at least about 42° C. Varying additional parameters, such as hybridization time, the concentration of detergent, e.g., sodium dodecyl sulfate (SDS), and the inclusion or exclusion of carrier DNA, are well known to those skilled in the art. Various levels of stringency are accomplished by combining these various conditions as needed. In certain embodiments, hybridization will occur at 30° C in 750 mM NaCl, 75 mM trisodium citrate, and 1% SDS. In certain embodiments, hybridization will occur at 37° C in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide, and 100 pg / ml denatured salmon sperm DNA (ssDNA). In certain embodiments, hybridization will occur at 42° C in 250 mM NaCl, 25 mM trisodium citrate, 1% SDS, 50% formamide, and 200 pg / ml ssDNA. Useful variations on these conditions will be readily apparent to those skilled in the art. For most applications, washing steps that follow hybridization will also vary in stringency. Wash stringency conditions can be defined by salt concentration and by temperature. As above, wash stringency can be increased by decreasing salt concentration or by increasing temperature. For example, stringent salt concentration for the wash steps can be less than about 30 mM NaCl and 3 mM trisodium citrate, e.g., less than about 15 mM NaCl and 1.5 mM trisodium citrate. Stringent temperature conditions for the wash steps will ordinarily include a temperature of at least about 25° C, of at least about 42° C, or of at least about 68° C. In certain embodiments, wash steps will occur at 25° C in 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. Incertain embodiments, wash steps will occur at 42° C. in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. In certain embodiments, wash steps will occur at 68° C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Additional variations on these conditions will be readily apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art and are described, for example, in Benton and Davis (Science 196: 180, 1977); Grunstein and Rogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wiley Interscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York.

[0083] By “substantially identical” or “substantially homologous” is meant a polypeptide or nucleic acid molecule exhibiting at least about 50% homologous or identical to a reference amino acid sequence (for example, any one of the amino acid sequences described herein) or nucleic acid sequence (for example, any one of the nucleic acid sequences described herein). In certain embodiments, such a sequence is at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, or at least about 100% homologous or identical to the sequence of the amino acid or nucleic acid used for comparison.

[0084] Sequence identity can be measured by using sequence analysis software (for example, Sequence Analysis Software Package of the Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705, BLAST, BESTFIT, GAP, or PILEUP / PRETTYBOX programs). Such software matches identical or similar sequences by assigning degrees of homology to various substitutions, deletions, and / or other modifications. Conservative substitutions typically include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine. In an exemplary approach to determining the degree of identity, a BLAST program may be used, with a probability score between e-3 and e-100 indicating a closely related sequence.

[0085] By “analog” is meant a structurally related polypeptide or nucleic acid molecule having the function of a reference polypeptide or nucleic acid molecule.

[0086] In the present application, the terms “variant transcript” or “variant transcript” or “chimeric transcripts” are used indifferently as synonyms. A “fusion or a variant or a chimeric” “transcript or sequence”, as per the present disclosure is defined as a transcript that aligns inpart with an exon sequence and in part with a non-exonic sequence, e.g., transposable element (TE) sequence. In some embodiments, the transcript has a normalized number of read greater than 2.10-6. The normalized number of reads is defined as the number of reads that cover the fusion divided by the library size of the sample.

[0087] Unless specifically stated or obvious from context, as used herein, the term “about” is to be understood as within a range of normal tolerance in the art, for example within 2 standard deviations of the mean. About can be understood as within 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.5%, 0.1%, 0.05%, or 0.01% of the stated value. Unless otherwise clear from context, all numerical values provided herein are modified by the term about.

[0088] A “transposable element” is to be understood as both class I (retrotransposons, including those containing LTRs, LINEs and SINEs) and class II (DNA transposons) endogenously part of the genome (i.e.: not from infection). This includes both autonomous and non-autonomous elements from both classes. According to the present disclosure the TE sequences can be for example selected from TE of class I, such as retrotransposons including Endogenous RetroVirus (ERVs), Long interspersed nuclear elements (LINEs) and short interspersed nuclear element (SINEs) and mammalian long terminal repeat transposon (MaLR), and TE of class II, such as DNA transposons endogenously part of the genome.

[0089] A reading frame is a way of dividing the sequence of nucleotides in a nucleic acid (DNA or RNA) molecule into a set of consecutive, non-overlapping triplets.

[0090] An open reading frame (ORF) is the part of a reading frame that has the ability to be translated into a peptide. An ORF is a continuous stretch of codons that contain a start codon (for example AUG) at a transcription starting site (TSS) and a stop codon (for example UAA, UAG or UGA). An ATG codon within the ORF (not necessarily the first) may indicate where translation starts. The transcription termination site is located after the ORF, beyond the translation stop codon. In eukaryotic genes with multiple exons, ORFs span intron / exon regions, which may be spliced together after transcription of the ORF to yield the final mRNA for protein translation.

[0091] A “canonical ORF” as herein intended is a protein coding sequence with specified reading frame within a mRNA sequence which is described or annotated in databases such as for example Ensembl genome / transcriptome / proteome database collection (typically HG19). Typically, a canonical ORF is the same as one of the exons in normal healthy cells.

[0092] A “non-canonical ORF” as herein intended is a protein coding sequence with specified reading frame within a mRNA sequence which is not described (i.e. unannotated) in genome databases such as for example in Ensembl genome / transcriptome / proteome database. Typically a non-canonical ORF means thus that the reading frame is shifted compared to the usual reading frame of exons in normal healthy cells. In some embodiments however, a non-canonical can be described in genome databases (such as Ensembl database), but the mRNA sequence represents minor species in normal cells. By minor species it is typically intended less that 5 % , notably less than 2 %, or preferentially less than 1 % species in normal cells.

[0093] An exon is any part of a gene that will encode a part of the final mature RNA produced by that gene after introns have been removed by RNA splicing. The term exon refers to both the DNA sequence within a gene and to the corresponding sequence in RNA transcripts. In RNA splicing, introns are removed and exons are covalently joined to one another as part of generating the mature messenger RNA. An exonic sequence as per the present applicant comprises at least a portion of one or more exon. Typically, the exonic sequence comprises at least a portion of one or 2 exons. Thus, intron sequences are considered non-exonic sequence.

[0094] The term “ polypeptide,” is used in the present specification to designate a series of residues, typically L-amino acids, connected one to the other, typically by peptide bonds between the a-amino and carboxyl groups of adjacent amino acids. The polypeptides or peptides can be a variety of lengths, either in their neutral (uncharged) forms or in forms which are salts, and either free of modifications such as glycosylation, side chain oxidation, or phosphorylation or containing these modifications, subject to the condition that the modification not destroy the biological activity of the polypeptides as herein described. Except is expressively mentioned the term “polypeptide” is used interchangeably with the term “protein”.

[0095] A “reference genome, or “representative genome” is a digital nucleic acid sequence data base, assembled by scientists as a representative example of species set of genes. As they are often assembled from the sequencing of DNA from a number of donors, reference genomes do not accurately represent the set of genes of any single individual (animal or person). Instead a reference provides a haploid mosaic of different DNA sequences from each donor.

[0096] As used herein, the term “antibody” means not only intact antibody molecules, but also fragments of antibody molecules that retain immunogen-binding ability. Such fragments are also well known in the art and are regularly employed both in vitro and in vivo. Accordingly,as used herein, the term “antibody” means not only intact immunoglobulin molecules but also the well-known active fragments F(ab’)2, and Fab.

[0097] F(ab’)2, and Fab fragments that lack the Fc fragment of intact antibody, clear more rapidly from the circulation, and may have less non-specific tissue binding of an intact antibody (Wahl et ah, J. Nucl. Med. 24:316-325 (1983).

[0098] As used herein, antibodies include whole native antibodies, bispecific antibodies; chimeric antibodies; Fab, Fab’, single chain V region fragments (scFv), fusion polypeptides, and unconventional antibodies.

[0099] In certain embodiments, an antibody is a glycoprotein comprising at least two heavy (H) chains and two light (L) chains inter-connected by disulfide bonds. Each heavy chain is comprised of a heavy chain variable region (abbreviated herein as VH) and a heavy chain constant (CH) region. The heavy chain constant region is comprised of three domains, CHI, CH2 and CH3. Each light chain is comprised of a light chain variable region (abbreviated herein as VL) and a light chain constant CL region. The light chain constant region is comprised of one domain, CL. The VH and VL regions can be further sub-divided into regions of hypervariability, termed complementarity determining regions (CDR), interspersed with regions that are more conserved, termed framework regions (FR). Each VH and VL is composed of three CDRs and four FRs arranged from amino-terminus to carboxy-terminus in the following order: FR1, CDR1, FR2, CDR2, FR3, CDR3, FR4.

[0100] The variable regions of the heavy and light chains contain a binding domain that interacts with an antigen. The constant regions of the antibodies may mediate the binding of the immunoglobulin to host tissues or factors, including various cells of the immune system ( e.g ., effector cells) and the first component (Cl q) of the classical complement system.

[0101] As used herein, “CDRs” are defined as the complementarity determining region amino acid sequences of an antibody which are the hypervariable regions of immunoglobulin heavy and light chains (See , e.g. , Rabat et al., Sequences of Proteins of Immunological Interest, 4thU. S. Department of Health and Human Services, National Institutes of Health (1987). Generally, antibodies comprise three heavy chain and three light chain CDRs or CDR regions in the variable region. CDRs provide the majority of contact residues for the binding of the antibody to the antigen or epitope. In certain embodiments, the CDRs regions are delineated using the Rabat system (Rabat, E. A., et al. (1991) Sequences of Proteins of ImmunologicalInterest, Fifth Edition, ET.S. Department of Health and Human Services, NIH Publication No. 91-3242).

[0102] As used herein, the term “single-chain variable fragment” or “scFv” is a fusion protein of the variable regions of the heavy (VH) and light chains (VL) of an immunoglobulin covalently linked to form a VH: :VL heterodimer. The VH and VL are either joined directly or joined by a peptide-encoding linker (e.g., 10, 15, 20, 25 amino acids), which connects the N- terminus of the VH with the C-terminus of the VL, or the C-terminus of the VH with the N- terminus of the VL. The linker is usually rich in glycine for flexibility, as well as serine or threonine for solubility. Despite removal of the constant regions and the introduction of a linker, scFv proteins retain the specificity of the original immunoglobulin. Single chain Fv polypeptide antibodies can be expressed from a nucleic acid including VH - and VL -encoding sequences as described by Huston, et al. (Proc. Nat. Acad. Sci. USA, 85:5879-5883, 1988). See, also , U.S. Patent Nos. 5,091,513, 5,132,405 and 4,956,778; and U.S. Patent Publication Nos. 20050196754 and 20050196754.

[0103] As used herein, the term “affinity” is meant a measure of binding strength. Affinity can depend on the closeness of stereochemical fit between antibody combining sites and antigen determinants, on the size of the area of contact between them, and / or on the distribution of charged and hydrophobic groups. As used herein, the term “affinity” also includes “avidity”, which refers to the strength of the antigen-antibody bond after formation of reversible complexes. Methods for calculating the affinity of an antibody for an antigen are known in the art, including, but not limited to, various antigen-binding experiments, e.g., functional assays (e.g., flow cytometry assay). Surface plasmon resonance assays such as BIACORE assays, and kinetic exclusion assays such as KINEXA assays

[0104] The term “chimeric antigen receptor” or “CAR” as used herein refers to a molecule comprising an extracellular antigen-binding domain that is fused to an intracellular signalling domain that is capable of activating or stimulating an immunoresponsive cell, and a transmembrane domain. In certain embodiments, the extracellular antigen-binding domain of a CAR comprises a scFv. The scFv can be derived from fusing the variable heavy and light regions of an antibody. Alternatively or additionally, the scFv may be derived from Fab’s (instead of from an antibody, e.g., obtained from Fab libraries). In certain embodiments, the scFv is fused to the transmembrane domain and then to the intracellular signaling domain. In certain embodiments, the CAR has a high binding affinity or avidity for the antigen.

[0105] The term “antigen-binding domain” as used herein refers to a domain capable of specifically binding a particular antigenic determinant or set of antigenic determinants present on a cell.

[0106] The term “immune cell” as herein intended typically encompasses T cells, Natural Killer T cells, CD4+ / CD8+ T cells, TILs / tumor derived CD8 T cells, central memory CD8+ T cells, Treg, MAIT, Y5 T cells, human embryonic stem cells, and pluripotent stem cells from which lymphoid cells may be differentiated.

[0107] By “isolated cell” is meant a cell that is separated from the molecular and / or cellular components that naturally accompany the cell.

[0108] The terms “isolated,” “purified,” or “biologically pure” refer to material that is free to varying degrees from components which normally accompany it as found in its native state. “Isolate” denotes a degree of separation from original source or surroundings. “Purify” denotes a degree of separation that is higher than isolation. A “purified” or “biologically pure” protein is sufficiently free of other materials such that any impurities do not materially affect the biological properties of the protein or cause other adverse consequences. That is, a nucleic acid or peptide is purified if it is substantially free of cellular material, viral material, or culture medium when produced by recombinant DNA techniques, or chemical precursors or other chemicals when chemically synthesized. Purity and homogeneity are typically determined using analytical chemistry techniques, for example, polyacrylamide gel electrophoresis or high performance liquid chromatography. The term “purified” can denote that a nucleic acid or protein gives rise to essentially one band in an electrophoretic gel. For a protein that can be subjected to modifications, for example, phosphorylation or glycosylation, different modifications may give rise to different isolated proteins, which can be separately purified.Method for selecting novel protein isoforms

[0109] The method for identifying a novel protein isoform, wherein the novel protein isoform comprises non-exonic amino acid sequence, can comprise the steps of:(a) aligning RNA sequences reads from a mammalian cell or tissue to a reference mammalian genome,(b) assembling transcripts from the genome-aligned RNA sequence reads of step (a), optionally using the combination of STAR and StringTie,(c) selecting variant transcripts from step (b) that (i) comprise one or more mutations compared to a reference mammalian genome, (ii) comprise exonic sequence, and (iii) comprise non- exonic sequence, optionally wherein the non-exonic sequence is a transposable element (TE),(d) optionally selecting variant transcripts from step (b) or step (c) that contain junctions between exonic sequence and non-exonic sequence and optionally are recurrently expressed in at least 1 % of subjects from a population of subjects suffering from a disease, optionally cancer, and (e) identifying the open reading frames (ORFs) of the variant transcripts from step (c) or (d) that are translated into protein, thereby identifying the novel protein isoform.

[0110] The method may further comprise selecting a neoantigentic peptide by further steps comprising: selecting a neoantigenic peptide of at least 8 amino acids, comprising a part of non- exonic amino acid sequence, and optionally selecting a peptide that binds to at least one Major Histocompatibility Complex (MHC) molecule of said subject.

[0111] Typically, a peptide comprising a part of non-exonic amino acid sequence is recognized as non-self by the immune system.

[0112] In any of the embodiments herein, the non-exonic sequence can be a TE, optionally intronic TE, optionally a short interspersed nuclear element (SINE), or a long interspersed nuclear element (LINE), or a long terminal repeat (LTRs). In some embodiments, the exonic sequence can be from a gene of phylostratum 1. In some embodiments, the exonic sequence is from a cancer driver gene, an oncogene and / or a tumor suppressor gene or their mutated variants.

[0113] Conceptually, cancer is a result of consecutive somatic mutation accumulation. Many studies have shown that both the gain of function in oncogenes and the loss of function in tumorsuppressor genes are required for the development of cancer from a normal cell. For a diploid organism, gain-of-function mutations are often dominant or semi-dominant, whereas loss-of- function mutations are usually recessive. Two-hit hypothesis of oncogenesis proposes that the development of cancer is initiated by the loss of both alleles of a tumor-suppressor gene.

[0114] Oncogenes (also named cancer genes) are genes whose action positively promotes cell proliferation or growth. The normal nonmutant versions are known as proto-oncogenes. The mutant versions are excessively or inappropriately active leading to tumor growth. Oncogenes can be identified in the Cancer Gene Marker Database (CGMD) (Pradeepkiran, J., Sainath, S., Kramthi Kumar, K. et al. CGMD:. Sci Rep 5, 12035 (2015) “An integrated database of cancer genes and markers”). Oncogenes (ONCs) can also be downloaded from Network of CancerGenes database (NCG 5.0) (A’ O, Dall'Olio GM, Mourikis TP, Ciccarelli FD, Nucleic Acids Res. 2016 Jan 4; 44(Dl):D992-9; “NCG 5.0: updates of a manually curated repository of cancer genes and associated properties from cancer mutational screenings”). Non-limitative examples of oncogenes include: L-MYC, LYL-1, LYT-10, LYT-10 / Cal, MAS, MDM-2, MLL, MOS, MTG8 / AML1, MYB, MYH11 / CBFB, NEU, N-MYC, OST, PAX-5, PBX1 / E2A, PIM-1, PRAD-1, RAF, RAR / PML, RAS-H, RAS-K, RAS-N, REL / NRG, RET, RH0M1, RH0M2, ROS, SKI, SIS, SET / CAN, SRC, TALI, TAL2, TAN-1, TIAM1, TSC2,and TRK.

[0115] Tumor suppressor genes (also named anti -oncogenes) represent the opposite side of cell growth control, normally acting to inhibit cell proliferation and tumor development. Thus tumor suppressor genes are genes that normally suppress cell division or growth. Loss of TSG function promotes uncontrolled cell division and tumor growth. Rb, a tumor suppressor gene that was identified by the genetic analysis of retinoblastoma an encoding a transcriptional regulatory protein, served as the prototype for the identification of additional tumor suppressor genes that contribute to the development of many different human cancers. Tumor suppressor genes are notably described in “Cooper GM. The Cell: A Molecular Approach. 2nd edition. Sunderland (MA): Sinauer Associates; 2000. Tumor Suppressor Genes”. Tumor-suppressor genes (TSGs) can also be downloaded from Tumor Suppressor Gene database (TSGene 2.0) (see for reference Zhao M, Kim P, Mitra R, Zhao J, Zhao Z; Nucleic Acids Res. 2016 Jan 4; 44(D1):D1O23-31; “TSGene 2.0: an updated literature-based knowledgebase for tumor suppressor genes”). In this context, non-limitative examples of tumor suppressor genes include: APC, BRCA1, BRCA2, DPC4, INK4, MADR2, NF1, NF2, p53, PTC, PTEN, Rb, RBI, VHL, WT1, BUB1, BUBR1, TGF-PRII, Axin, DPC4, p300, PPARy, pl 6, DPC4, PTEN, and hSNF5.

[0116] Oncogenes, tumor suppressor genes or “double agent” genes (with both oncogenic and tumor-suppressor functions) can be systematically identified through database search and text mining. Indeed, information on oncogenes or tumor suppressor genes can typically be found in Ensembl database (but see also Shen L, Shi Q, Wang W. Double agents: genes with both oncogenic and tumor-suppressor functions. Oncogenesis. 2018;7(3):25. Published 2018 Mar 13). Double agent genes may be identified as genes overlapped between the two above mentioned databases (see also Shen et al., Oncogenesis 2018 above).

[0117] Without being bound by any theory, the selection of fusion wherein the exonic sequence is from a cancer driver gene, an oncogene, and / or a tumor suppressor gene is of high relevance for the reason below:

[0118] Insertion of non-exonic sequence, e.g. TE sequence, in oncogenes can alter their oncogenic activity. Insertion of non-exonic sequence, e.g. TE sequence, in oncogene active domains could therefore result in constitutive activity of the oncogenes, similar to driver mutations. These fusions giving chimeric oncogenes could thus represent a new family of oncogenic proteins. If this is the case, targeting the activity of these new "fusion oncogenes" with small molecule antagonists could represent a potential therapeutic approach for cancer where these chimeric oncogenes are expressed.

[0119] Insertion of non-exonic sequence, e.g. TE sequence in tumor suppressors could inactivate their suppressor functions, leading typically to a loss of function (for example through introduction of stop codons, changes in ORF or disruptive amino acid stretches), thereby contributing to the oncogenic process.

[0120] Insertions of non-exonic sequence, e.g. TE sequence, implicating cancer driver genes would be excellent targets for adoptive cell therapies, antibodies, ADCs, T cell engagers, etc. If they are involved in oncogenesis, variant oncogenes are expected to be more specific for cancer cells, and thus to reduce the development of resistances (because of the oncogenic activity of the target).

[0121] In one embodiment, the non-exonic sequence, e.g. TE sequence, is located in 5’ end of the variant transcript sequence (it is also said that the non-exonic sequence is the donor sequence) and the exonic sequence is located in 3’ end of the variant transcript sequence with respect to the junction (the exon sequence is thus called an acceptor sequence). The expression “is located in 5’ end of the variant transcript sequence” means that the element is located upstream of the junction in the variant transcript sequence. The expression “is located in 3’ end of the variant transcript sequence” means that the element is located downstream of the junction in the variant transcript sequence.

[0122] In a particular embodiment, the non-exonic sequence, e.g. TE sequence, is located in 5’ end of the variant transcript sequence and the exonic sequence is located in 3’ end of the variant transcript sequence, and the part of the ORF of said variant transcript sequence, which encodes the novel protein isoform, overlaps the junction. In this case, the ORF can be canonical or non- canonical. For example, the ORF may be frameshifted relative to the canonical ORF and therefore produced completely different non-exonic amino acid sequence. It is understood that the ORF may comprise the junction but the non-exonic amino acid sequence need not comprisethe junction. In some embodiments, where the neoantigenic peptide comprises the junction, it comprises both non-exonic amino acid sequence and exonic amino acid sequence.

[0123] The expression “the part of the ORF is overlapping or overlaps the junction between the non-exonic sequence, e.g. TE sequence, and the exonic sequence”, means that said junction is contained in the part of the ORF which encodes said neoantigenic peptide.

[0124] In embodiments wherein (i) the part of the ORF encoding the neoantigenic peptide is overlapping the junction between the non-exonic sequence, e.g. TE sequence, and the exonic sequence, and (ii) the non-exonic sequence, e.g. TE sequence, and the exonic sequence are respectively in 5 ’ end and 3 ’ end of the variant transcript sequence, said part of the ORF typically encodes a neoantigenic peptide of at least 8 amino acids, including at least between 1 to 6 amino acids, notably 2 to 6 from the non-exonic sequence, e.g. TE sequence, and at least between 1 and 6, notably 2 to 6 amino acids from the exonic sequence.

[0125] In another embodiment wherein the non-exonic sequence, e.g. TE sequence, is located in 5’ end of the variant transcript sequence and the exonic sequence is located in 3’ end of the variant transcript sequence, the part of ORF which encodes said neoantigenic peptide, is downstream of the junction and the ORF is thus non-canonical.

[0126] The expression “the part of the ORF is downstream of the junction” means that the part of the ORF encoding the neoantigenic peptide is not overlapping the junction, but it is contained in the 3 ’end part of said variant transcript sequence with respect to the junction. In this embodiment, as the 3’ end part with respect to the junction, the obtained peptide is encoded by exonic sequence as part of a non-canonical ORF. This non-canonical ORF is also referred to herein as non-exonic amino acid sequence, because the encoded amino acid sequence is not part of a canonical exon. Thus, in the particular embodiment wherein the exonic sequence is located in 3’ end of the variant transcript sequence with respect to the junction, and wherein the part of the ORF which encodes the neoantigenic peptide is downstream of the junction with a non-canonical reading frame, the part of the ORF of the variant transcript sequence encodes a neoantigenic peptide including 0 amino acid from the non-exonic sequence, e.g. TE sequence, and at least 8 amino acids from the exonic sequence.

[0127] In another embodiment, the non-exonic sequence, e.g. TE sequence, is located in 3’ end of the variant transcript sequence and the exonic sequence is located in 5’ end of the variant transcript sequence with respect to the junction.

[0128] In some embodiments, the non-exonic sequence, e.g. TE sequence is located in 3’ end of the variant transcript sequence and the exonic sequence is located in 5’ end of the variant transcript sequence and the part of the ORF of said variant transcript sequence, which encodes a neoantigenic peptide, is overlapping the junction between the non-exonic sequence, e.g. TE sequence, and the exonic sequence. In this case, the ORF can also be canonical or non- canonical. The obtained peptide is encoded by both non-exonic sequence, e.g. TE sequence, and exonic sequence.

[0129] In the particular embodiment wherein the part of the ORF encoding the neoantigenic peptide, is overlapping the junction between the exonic sequence and the non-exonic sequence, e.g. TE sequence, and wherein the exonic sequence and the non-exonic sequence, e.g. TE sequence, are respectively in 5’ end and 3 ’end of the variant transcript sequence, said part of the ORF encodes a neoantigenic peptide of at least 8 amino acids, including at least between 1 to 6, notably 2 to 6 amino acids from the non-exonic sequence, e.g. TE sequence, and at least between 1 and 6, notably 2 to 6 amino acids from the exonic sequence.

[0130] In still another embodiment, the non-exonic sequence, e.g. TE sequence, is located in 3’ end of the variant transcript sequence, the exonic sequence is located in 5’ end of the variant transcript sequence, and the part of the ORF which encodes a neoantigenic peptide, is downstream of the junction between the exonic sequence and the non-exonic sequence, e.g. TE sequence. Optionally, the peptide sequence which is thus encoded by the pure non-exonic sequence, e.g. TE sequence, is non-canonical.

[0131] In this embodiment, as the 3’ end part with respect to the junction is the non-exonic sequence, e.g. TE sequence, the part of the ORF encoding the neoantigenic peptide is therefore encoded by the non-exonic sequence, e.g. TE sequence. Thus, the part of the ORF encodes a neoantigenic peptide including no amino acid from the exonic sequence and at least 8 amino acids from the non-exonic sequence, e.g. TE sequence. In the particular embodiment wherein the non-exonic sequence, e.g. TE sequence, is located in 3’ end of the variant transcript sequence with respect to the junction, and the part of the ORF which encodes the neo antigenic peptide is downstream the junction, the part of the ORF of the variant transcript sequence encodes a neoantigenic peptide including 0 amino acid from the exonic sequence, and at least 8 amino acids from the non-exonic sequence, e.g. TE sequence.

[0132] A neoantigenic peptide is a peptide that arises from somatic alterations (classically mutations in the DNA sequence), is recognized as different from self, and is presented byantigen-presenting cells (APC), such as dendritic cells (DC) and tumor cells themselves. Crosspresentation plays an important role as the APC is able to translocate exogenous antigens from the phagosome into the cytosol for proteolytic cleavage into the major histocompatibility complex I (MHC I) epitopes by the proteasome.

[0133] In the present disclosure the alteration corresponds to the variant transcript that comprise a non-exonic sequence, e.g. TE sequence, and an exonic sequence. This may arise from somatic (e.g., in the tumor clone) transposition. It may also arise not from de novo transposition but from tumor specific transcriptional de-repression such that non-exonic sequence, e.g. TE sequence, and nearby gene are co-transcribed.

[0134] A neoantigenic peptide according to the present disclosure may be completely absent from normal healthy samples (i.e., not expressed in normal healthy samples) and thus be specific to tumor samples. Alternatively, it may be expressed at low levels in normal cells and / or disproportionately expressed on disease, e.g. tumor, samples as compared to normal (healthy) samples.

[0135] It can also be selectively expressed by the cell lineage from which the cancer evolved.

[0136] Cancer or tumor samples according to the present disclosure can be isolated from any solid tumor or non-solid tumor of any tissues or organs, for example, breast cancer, lung cancer and / or melanoma. In some embodiments cancer samples are from Acute Myeloid Leukemia, Adrenocortical Carcinoma, Bladder Urothelial Carcinoma, Breast Ductal Carcinoma, Breast Lobular Carcinoma, Cervical Carcinoma, Cholangiocarcinoma, Colorectal Adenocarcinoma, Esophageal Carcinoma, Gastric Adenocarcinoma, Glioblastoma Multiforme, Head and Neck Squamous Cell Carcinoma, Hepatocellular Carcinoma, Kidney Chromophobe Carcinoma, Kidney Clear Cell Carcinoma, Kidney Papillary Cell Carcinoma, Lower Grade Glioma, Lung Adenocarcinoma, Lung Squamous Cell Carcinoma, Mesothelioma, Ovarian Serous Adenocarcinoma, Pancreatic Ductal Adenocarcinoma, Paraganglioma & Pheochromocytoma, Prostate Adenocarcinoma, Sarcoma, Skin Cutaneous Melanoma, Testicular Germ Cell Cancer, Thymoma, Thyroid Papillary Carcinoma, Uterine Carcinosarcoma, Uterine Corpus Endometrioid Carcinoma or Uveal Melanoma samples. In a particular embodiment, cancer samples are from lung cancer samples, notably from LU AD samples.

[0137] Such a method for identifying the novel protein isoform may further comprise the step of selecting variant transcripts present in mRNA subpopulations of a mammalian cell or tissue identified as directly bound to ribosomes, optionally using RiboSeq sequencing, RiboseQCand / or ORFquant; and optionally comprising the step of selecting the ORFs of step (e) that are present in peptide fragments identified through mass spectrometry proteomics. In such methods, the non-exonic region of the reference mammalian genome can be a TE, optionally intronic TE, optionally a short interspersed nuclear element (SINE), or a long interspersed nuclear element (LINE), or a long terminal repeat (LTRs). In such methods, the exonic sequence can be from a gene of phylostratum 1.

[0138] In one embodiment, the mRNA sequences can be mapped against a corresponding reference genome, with an adapted software, such as for example: Spliced Transcripts Alignment to a Reference -i.e.: STAR - see Dobin, Alexander et al. “STAR: ultrafast universal RNA-seq aligner.” Bioinformatics (Oxford, England) vol. 29,1 (2013): 15-21), TopHat2 (Kim, Daehwan et al. “TopHat2: accurate alignment of transcriptomes in the presence of insertions, deletions and gene fusions.” Genome biology vol. 14,4 R36. 25 Apr. 2013, doi: 10.1186 / gb- 2013-14-4-r36) or HISAT (Kim, Daehwan et al. “HISAT: a fast spliced aligner with low memory requirements.” Nature methods vol. 12,4 (2015): 357-60. doi: 10.1038 / nmeth.3317). STAR is a standalone software that uses sequential maximum mappable seed search followed by seed clustering and stitching to align RNA-seq reads. It is able to detect canonical junctions, non-canonical splices, and fusion / chimeric transcripts.

[0139] Alternatively, or in addition, the mRNA reads are assembled using StringTie (Pertea et al., Nat Biotechnol. 2015 Mar; 33(3): 290-295). StringTie is a fast and highly efficient assembler of RNA-Seq alignments into potential transcripts. It uses a novel network flow algorithm as well as an optional de novo assembly step to assemble and quantitate full-length transcripts representing multiple splice variants for each gene locus. Its input can include not only alignments of short reads that can also be used by other transcript assemblers, but also alignments of longer sequences that have been assembled from those reads. In order to identify differentially expressed genes between experiment’, StringTie's output can be processed by specialized software like Ballgown, Cuffdiff or other programs (DESeq2, edgeR, etc.).

[0140] In a particular embodiment, the method involves use of Riboseq alignment, a method that permits alignment to sequences from individual ribosome footprints. See, e.g., Calviello et al., Trends in Genetics, vol. 33(10), 728-744, 2017. In such methods, the exact codon being translated in a purified ribosome-RNA complex can be determined. Through the sequencing of hundreds of millions of ribosome footprints, a single Ribo-seq experiment can therefore produce a detailed and accurate representation of a given sample’s translated RNAs. The method may optionally involve use of open reading frame (ORF) detection using a method suchas ORF quant, a method that detects and quantifies ORF translation on complex transcriptomes using the Riboseq data. See, e.g., Calviello et al., Nat Struct Mol Biol 27, 717-725 (2020).

[0141] According to the present disclosure, the mRNA sequences can come from all types of cancer cell or tumor cell sample(s). The tumor may be a solid or a non-solid tumor. In particular, the mRNA sequences come from any tissues or organs affected by a cancer or tumor as previously defined, for example from breast cancer, lung cancer and / or melanoma. In a particular embodiment, mRNA sequences are from TCGA or CCLE or LU AD samples.

[0142] Typically, as per the present disclosure, the variant transcript sequences are shared in more than 1%; notably more than 5%, more than 10%, more than 15%, more than 20% or even more than 25 % of the cancer samples. In other words, a variant transcript sequence as per the present disclosure is shared in cancer samples from more than 1%; notably more than 5%, more than 10%, more than 15%, more than 20% or even more than 25 % of the subjects suffering from a cancer. The variant transcript sequence may thus be specific for a cancer type of shared between several cancers.

[0143] According to the present disclosure, the variant transcript sequences are expressed at higher levels in cells associated with disease, e.g., tumor cells, compared to normal healthy cells. In some embodiments, the variant transcript sequence is expressed in cells associated with disease, e.g., cancer cells, and not in healthy cells, in particular not in thymus healthy cells. Such variant transcript may be called disease specific fusion, or tumor specific fusion as per the present disclosure. Variant transcripts that are expressed at higher level(s) in tumor cells as compared to normal cell, typically that are disproportionally expressed in cancers cells as compared to normal cells as defined above may be called tumor associated variant transcripts (TAF) as per the present disclosure. Tumor associated variant transcripts may be selected according to the present application if they are present in more than 10% of the tumor samples and in less than 20% of the normal samples.

[0144] In some embodiments, the method further comprises a step of determining, optionally in silico or using in vitro techniques (see notably the example for illustration), the binding affinity of the neoantigenic peptide with at least one MHC molecule of a human subject, a subject suffering from disease.

[0145] MHC class I proteins form a functional receptor on most nucleated cells of the body. There are 3 major MHC class I genes in HLA: HLA-A, HLA-B, HLA-C and three minor genes HLA-E, HLA-F and HLA-G. P2-microglobulin binds with major and minor gene subunits toproduce a heterodimer. MHC molecules of class I consist of a heavy chain and a light chain and are capable of binding a peptide of about 8 to 11 amino acids, but usually 8 or 9 amino acids, if this peptide has suitable binding motifs, and presenting it to cytotoxic T- lymphocytes. The binding of the peptide is stabilized at its two ends by contacts between atoms in the main chain of the peptide and invariant sites in the peptide-binding groove of all MHC class I molecules. There are invariant sites at both ends of the groove which bind the amino and carboxy termini of the peptide. Variations in peptide length are accommodated by a kinking in the peptide backbone, often at proline or glycine residues that allow the required flexibility. The peptide bound by the MHC molecules of class I usually originates from an endogenous protein antigen. As an example, the heavy chain of the MHC molecules of class I is typically an HLA-A, HLA- B or HLA-C monomer, and the light chain is P-2-microglobulin, in humans.

[0146] There are 3 major and 2 minor MHC class II proteins encoded by the HL A. The genes of the class II combine to form heterodimeric (a[3) protein receptors that are typically expressed on the surface of antigen-presenting cells. The peptide bound by the MHC molecules of class II usually originates from an extracellular or exogenous protein antigen. As an example, the a -chain and the P-chain are in particular HLA-DR, HLA-DQ and HLA-DP monomers, in humans. MHC class II molecules are capable of binding a peptide of about 8 to 20 amino acids, notably from 10 to 25 or from 13 to 25 if this peptide has suitable binding motifs and presenting it to T-helper cells. These peptides lie in an extended conformation along the MHC II peptide- binding groove which (unlike the MHC class I peptide-binding groove) is open at both ends. The peptide is held in place mainly by main-chain atom contacts with conserved residues that line the peptide-binding groove.

[0147] When the method is carried out on human samples, the method may comprise a step of determining the patient’s class I or class I Major Histocompatibility Complex (MHC, aka human leukocyte antigen (HLA) alleles). It is to be noticed that as MHC alleles for laboratory mice are generally known such that this step may not be necessary in that particular context. In the present application, “MHC molecule” refers to at least one MHC class I molecule or at least one MHC Class II molecule.

[0148] A MHC allele database is carried out by analyzing known sequences of MHC I and MHC II and determining allelic variability for each domain. This can be typically determined in silico using appropriate software algorithms well-known in the field. Several tools have been developed to obtain HLA allele information from genome-wide sequencing data (whole-exome, whole-genome, and RNA sequencing data), including OptiType, Polysolver, PHLAT,HLAreporter, HLAforest, HLAminer, and seq2HLA (see Kiyotani K et al., Immunopharmacogenomics towards personalized cancer immunotherapy targeting neoantigens; Cancer Science 2018; 109:542-549). For example, the seq2hla tool (see Boegel S, Lower M, Schafer M, et al. HLA typing from RNA-Seq sequence reads. Genome Med. 2012;4: 102), which is well designed to perform the method as herein disclosed is an in silico method written in python and R, which takes standard RNA-Seq sequence reads in fastq format as input, uses a bowtie index (Langmead B, et al., Ultrafast and memory-efficient alignment of short DNA sequences to the human genome. Genome Biol. 2009, 10: R25-10.1186 / gb-2009- 10-3-r25) comprising all HLA alleles and outputs the most likely HLA class I and class II genotypes (in 4 digit resolution), a p-value for each call, and the expression of each class.

[0149] Typically, the non-exonic amino acid sequences are extracted in silico. Such non-exonic amino acid sequences include amino acid sequences encoded by a part of an ORF comprising a junction between a non-exonic sequence, e.g. TE sequence, and an exonic sequence. Such non-exonic amino acid sequences also include amino acid sequence resulting from a non- canonical ORF, e.g. resulting from a frameshift in ORF compared to canonical ORF. The affinity of all possible peptides encoded by each sequence for each MHC allele from the patient (or mouse) can be for example determined in silico using computational methods to predict peptide binding-affinity to HLA molecules. Indeed, accurate prediction approaches are based on artificial neural networks with predicted IC50. For example, NetMHCpan software which has been modified from NetMHC to predict peptides binding to alleles for which no ligands have been reported, is well appropriate to implement the method as herein disclosed (Lundegaard C et al., NetMHC-3.0: accurate web accessible predictions of human, mouse and monkey MHC class I affinities for peptides of length 8-11; Nucleic Acids Res. 2008;36:W509- W512; Nielsen M et al. NetMHCpan, a method for quantitative predictions of peptide binding to any HLA-A and -B locus protein of known sequence. PLoS One. 2007;2:e796, but see also Kiyotani K et al., Immunopharmacogenomics towards personalized cancerimmunotherapy targeting neoantigens; Cancer Science 2018; 109:542-549 and Yarchoan M et al., Nat rev. cancer 2017; 17(4):209-222). NetMHCpan software predicts binding of peptides to any MHC molecule of known sequence using artificial neural networks (ANNs). The method is trained on a combination of more than 180,000 quantitative binding data and MS derived MHC eluted ligands. The binding affinity data covers 172 MHC molecules from human (HLA-A, B, C, E), mouse (H-2), cattle (BoLA), primates (Patr, Mamu, Gogo) and swine (SLA). The MS eluted ligand data covers 55 HLA and mouse alleles.

[0150] In example embodiments, neoantigenic peptides encoded by variant transcripts as above described, comprising some non-exonic amino acid sequence and optionally some exonic amino acid sequence, are selected as neoantigenic peptides. In some embodiments, the selected neoantigenic peptides have a predicted or actual Kd affinity for MHC alleles of less than 10'4. 10-5, |Q-6, JO-7M or less than 500 nM, notably less than 50 nM.

[0151] As above mentioned, affinity of the selected peptide for MHC alleles can be determined in silico using appropriate software such as netMHCpan. Thus, in some embodiments, neoantigenic peptides bind MHC class I with a binding affinity of less than 2% percentile rank score predicted by NetMHCpan 4.0. In other embodiments, the neoantigenic peptides bind MHC class II with a binding affinity of less than 10% percentile rank score predicted by NetMHCpanll 3.2.

[0152] Affinity can also (alternatively or in addition) be estimated in vitro, for example using MHC tetramer formation assay as described in the results included therein (see example 2, point 2.1 and 2.2.2). Commercial assays for example from ImmunAware® can typically be used by the skilled person (EasYmers® kits are from ImmunAware® are notably used according to their training guide). Typically, binding affinity is determined as a percentage of binding to a positive control. Generally, peptides showing a percentage of binding of at least 30 %, notably at least 40% or even at least 50 % of the positive control are selected. Typically, the neoantigenic peptide as per the present disclosure, and typically obtainable as per the present method, binds at least one HLA / MHC molecule with an affinity sufficient for the peptide to be presented on the surface of a cell as an antigen. Generally, the neoantigenic peptide has an IC50 affinity of less than 10'4. or 10'5, or 10'6, or 10'7or less than 500 nM, at least less than 250nM, at least less than 200 nM, at least less than 150 nM, at least less than 100 nM, at least less than 50 nM or less for at least one HLA / MHC molecule (lower numbers indicating greater binding affinity), typically a molecule of said subject suffering from a disease, e.g., cancer.

[0153] Further optional steps according to the present method may thus independently include:

[0154] a step of exclusion of variant transcripts or predicted peptides expressed at high levels or high frequency on healthy cells. An alignment of the variant transcript sequence against the RNAseq data of healthy cells, typically allows determining the relative amount of variant transcript sequence(s) present in healthy cells; In one embodiment, variant transcripts or predicted peptides expressed on healthy cells are discarded.

[0155] a step to confirm that a neoantigenic peptide is not expressed in healthy cells of the subject. This step can be carried out using typically the Basic local alignment search tool (BLAST) and performing alignment of the sequence of the neoantigenic peptide against the proteome of healthy cells; Preferably, peptides that align against the proteome of normal healthy cells (for example using BLAST) are discarded.

[0156] a step to confirm that the variant transcript or predicted peptide is expressed in cells associated with disease, e.g., cancer cells, of the subject. The presence of the selected variant transcript sequence in cancer cells can be checked typically by RT-PCR in mRNA extracted from cancer cell sample.Neoantigenic peptides

[0157] The present disclosure also relates to an isolated neoantigenic peptide comprising at least 8, 9, 10, 11, or 12 amino acids, encoded by a portion of an open reading frame (ORF) from a variant transcript that is a human mRNA sequence comprising a non-exonic sequence, e.g. TE sequence, and an exonic sequence. The peptide may be 8-9, 8-10, 8-11, 12-25, 13-25, 12- 20, or 13-20 amino acids in length. Although the ORF overlaps a junction between a non-exonic sequence, e.g. TE sequence, and an exonic sequence, it is understood that the neoantigenic peptide itself may not comprise the junction.

[0158] In some embodiments, the variant transcript is from a cancer cell. In some embodiments, the variant transcript is associated with cancer. In example embodiments the neoantigenic peptide comprising at least 8, 9, 10, 11 or 12 amino acids is encoded by a part of an open reading frame (ORF) of any of the variant transcript sequences of any one of the nucleotide sequences of Table 1A, preferably a peptide with one or more of the neoantigenic peptide characteristics described above.

[0159] The peptide may be 8-9, 8-10, 8-11, 12-25, 13-25, 12-20, or 13-20 amino acids in length and fulfills one or more of the neoantigenic peptide characteristics described above. The N- terminus of the peptide of at least 8 amino acids may be encoded by the triplet codon starting at any of nucleotide positions 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, and higher (it being understood that the disclosure contemplates a start position that is any of the integers between 1 and 8000 without having to list every number between 1 and 8000).

[0160] A peptide as above defined is typically obtainable according to the method of the present disclosure and thus encompasses one or more of the characteristics as previously described. In particular a neoantigenic peptide as per the present disclosure may exhibit one or a combination of the following further characteristics:

[0161] It binds or specifically binds MHC class I of a subject and is 8 to 11 amino acids, notably 8, 9, 10, or 11 amino acids. Typically the neoantigenic peptide is 8 or 9 amino acids long, and binds to at least one MHC class I molecule of the subject; or alternatively, it binds to at least one MHC class II molecule of said subject and contains from 12 to 25 amino acids, notably is 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 amino acids long.

[0162] It binds at least one HLA / MHC molecule of said subject suffering from a disease, e.g., cancer with an affinity sufficient for the peptide to be presented on the surface of a cell as an antigen. Typically the neoantigenic peptide has an IC50 of less than 10'4. or 10'5, or 10'6, or 10’7or less than 500 nM, at least less than 250nM, at least less than 200 nM, at least less than 150 nM, at least less than 100 nM, at least less than 50 nM or less (lower numbers indicating greater binding affinity).

[0163] It does not induce a significant autoimmune response and / or invoke immunological tolerance when administered to a subject.

[0164] It is expressed at higher levels in tumor samples compared to normal healthy samples. Typically, as per the present disclosure, a variant transcript may be selected if it is present in more than 10 % of the tumor samples and in less than 20 % of the normal samples. In some embodiments, the neoantigenic is more specifically a tumor specific antigen (TSA), i.e.: it is only expressed in cancer sample and not in normal samples, or is expressed at relatively low levels in normal samples (e.g. the expressed mRNA sequences represent minor species in normal cells from normal samples).

[0165] It comprises the junction between the TE sequence and the exonic sequence, in other words it is encoded by a part of a TE sequence and a part of an exonic sequence, the ORF being either canonical or non-canonical or

[0166] It is encoded by a non-canonical ORF of an exonic sequence or

[0167] It is encoded by the TE sequence, optionally in a non-canonical ORF

[0168] A neoantigenic peptide may first be validated by RT transcription analysis of variant transcripts sequence in tumors cell from a subject. Typically also, immunization with a neoantigenic peptide as per the present disclosure elicits a T cell response

[0169] In a particular embodiment, the present disclosure encompasses a NSCLC neoantigenic peptide comprising at least 8 amino acids of any one of the amino acid sequences of Table 1A. Typically, said neoantigenic peptides of any of the amino acid sequences of Table 1A binds to HLA-A02 with an affinity sufficient for the peptide to be presented on the surface of cells as an antigen. Affinity for MHC alleles can be determined by known techniques in the field and notably in silico or in vitro as exemplified above;

[0170] In a particular embodiment, a neoantigenic peptide as per the present disclosure binds to a MHC molecule present in at least 1 %, 5 %, 10 %, 15 %, 20 %, 25% or more of subjects. Notably, a neoantigenic peptide as herein disclosed is expressed in at least 1 %, 5 %, 10 %, 15 %, 20 %, 25% of subjects from a population of subjects suffering from cancer.

[0171] More particularly, a neoantigenic peptide of the present disclosure is capable of eliciting an immune response against a tumor present in at least 1 %, 5 %, 10 %, 15 %, 20%, or 25 % of the subjects in the population of subjects suffering from cancer.

[0172] As previously defined, cancer may affect any one of the following tissues or organs: breast; liver; kidney; heart, mediastinum, pleura; floor of mouth; lip; salivary glands; tongue; gums; oral cavity; palate; tonsil; larynx; trachea; bronchus, lung; pharynx, hypopharynx, oropharynx, nasopharynx; esophagus; digestive organs such as stomach, intrahepatic bile ducts, biliary tract, pancreas, small intestine, colon; rectum; urinary organs such as bladder, gallbladder, ureter; rectosigmoid junction; anus, anal canal; skin; bone; joints, articular cartilage of limbs; eye and adnexa; brain; peripheral nerves, autonomic nervous system; spinal cord, cranial nerves, meninges; and various parts of the central nervous system; connective, subcutaneous and other soft tissues; retroperitoneum, peritoneum; adrenal gland; thyroid gland; endocrine glands and related structures; female genital organs such as ovary, uterus, cervix uteri; corpus uteri, vagina, vulva; male genital organs such as penis, testis and prostate gland; hematopoietic and reticuloendothelial systems; blood; lymph nodes; thymus. For example, the tumors or cancers as per the present application includes leukemias, seminomas, melanomas, teratomas, lymphomas, neuroblastomas, gliomas, rectal cancer, endometrial cancer, kidney cancer, adrenal cancer, thyroid cancer, blood cancer, skin cancer, cancer of the brain, cervical cancer, intestinal cancer, liver cancer, colon cancer, stomach cancer, intestine cancer, head andneck cancer, gastrointestinal cancer, lymph node cancer, esophagus cancer, colorectal cancer, pancreas cancer, ear, nose and throat (ENT) cancer, breast cancer, prostate cancer, cancer of the uterus, ovarian cancer and lung cancer and the metastases thereof. Examples thereof are lung carcinomas, mamma carcinomas, prostate carcinomas, colon carcinomas, renal cell carcinomas, cervical carcinomas, or metastases of the cancer types or tumors described above. The term cancer according to the present disclosure also comprises cancer metastases and relapse of cancer.

[0173] Typically a neoantigenic peptide as per the present disclosure does not induce a significant autoimmune response and / or invoke immunological tolerance when administered to a subject. Tolerating mechanisms involve clonal deletion, ignorance, anergy, or suppression in the host w the reduction in the number of high-affinity self-reactive T cells.

[0174] The neoantigenic peptide can also be modified by extending or decreasing the compound's amino acid sequence, e.g., by the addition or deletion of amino acids. The peptides can also be modified by altering the order or composition of certain residues, it being readily appreciated that certain amino acid residues essential for biological activity, e.g., those at critical contact sites or conserved residues, may generally not be altered without an adverse effect on biological activity. The non-critical amino acids need not be limited to those naturally occurring in proteins, such as L-a-amino acids, or their D-isomers, but may include non-natural amino acids as well, such as P-y-5-amino acids, as well as many derivatives of L-a-amino acids.

[0175] Typically, a series of peptides with single amino acid substitutions are employed to determine the effect of electrostatic charge, hydrophobicity, etc. on binding. For instance, a series of positively charged (e.g., Lys or Arg) or negatively charged (e.g., Glu) amino acid substitutions are made along the length of the peptide revealing different patterns of sensitivity towards various MHC molecules and T cell receptors. In addition, multiple substitutions using small, relatively neutral moieties such as Ala, Gly, Pro, or similar residues may be employed. The substitutions may be homo-oligomers or hetero-oligomers. The number and types of residues which are substituted or added depend on the spacing necessary between essential contact points and certain functional attributes which are sought (e.g., hydrophobicity versus hydrophilicity). Increased binding affinity for an MHC molecule or T cell receptor may also be achieved by such substitutions, compared to the affinity of the parent peptide. In any event, such substitutions should employ amino acid residues or other molecular fragments chosen to avoid, for example, steric and charge interference which might disrupt binding.

[0176] Amino acid substitutions are typically of single residues. Substitutions, deletions, insertions or any combination thereof may be combined to arrive at a final peptide. Substitutional variants are those in which at least one residue of a peptide has been removed and a different residue inserted in its place. Such substitutions are generally made in accordance with the following Table A when it is desired to finely modulate the characteristics of the peptide.TABLE A

[0177] Substantial changes in function (e.g., affinity for MHC molecules or T cell receptors) are made by selecting substitutions that are less conservative than those in above Table, i.e., selecting residues that differ more significantly in their effect on maintaining (a) the structure of the peptide backbone in the area of the substitution, for example as a sheet or helical conformation, (b) the charge or hydrophobicity of the molecule at the target site or (c) the bulk of the side chain. The substitutions which in general are expected to produce the greatest changes in peptide properties will be those in which (a) hydrophilic residue, e.g. seryl, is substituted for (or by) a hydrophobic residue, e.g. leucyl, isoleucyl, phenylalanyl, valyl or alanyl; (b) a residue having an electropositive side chain, e.g., lysl, arginyl, or histidyl, is substituted for (or by) an electronegative residue, e.g. glutamyl or aspartyl; or (c) a residue having a bulky side chain, e.g. phenylalanine, is substituted for (or by) one not having a side chain, e.g., glycine.

[0178] The peptides and polypeptides may also comprise isosteres of two or more residues in the neoantigenic peptide or polypeptides. An isostere as defined here is a sequence of two or more residues that can be substituted for a second sequence because the steric conformation of the first sequence fits a binding site specific for the second sequence. The term specifically includes peptide backbone modifications well known to those skilled in the art. Such modifications include modifications of the amide nitrogen, the a-carbon, amide carbonyl, complete replacement of the amide bond, extensions, deletions or backbone crosslinks. See, generally, Spatola, Chemistry and Biochemistry of Amino Acids, Peptides and Proteins, Vol. VII (Weinstein ed., 1983 ).

[0179] In addition, the neoantigenic peptide may be conjugated to a carrier protein, a ligand, or an antibody. Half-life of the peptide may be improved by PEGylation, glycosylation, polysialylation, HESylation, recombinant PEG mimetics, Fc fusion, albumin fusion, nanoparticle attachment, nanoparticulate encapsulation, cholesterol fusion, iron fusion, or acylation.

[0180] Modifications of peptides and polypeptides with various amino acid mimetics or unnatural amino acids are particularly useful in increasing the stability of the peptide and polypeptide in vivo. Stability can be assayed in a number of ways. For instance, peptidases and various biological media, such as human plasma and serum, have been used to test stability. See,e.g., Verhoef et al., Eur. J. Drug Metab Pharmacokin. 11 :291-302 (1986 ). Half life of the peptides of the present disclosure is conveniently determined using a 25% human serum (v / v) assay. The protocol is generally as follows. Pooled human serum (Type AB, non-heat inactivated) is delipidated by centrifugation before use. The serum is then diluted to 25% with RPMI tissue culture media and used to test peptide stability. At predetermined time intervals a small amount of reaction solution is removed and added to either 6% aqueous trichloracetic acid or ethanol. The cloudy reaction sample is cooled (4°C) for 15 minutes and then spun to pellet the precipitated serum proteins. The presence of the peptides is then determined by reversed-phase HPLC using stability-specific chromatography conditions.

[0181] The peptides and polypeptides may be modified to provide desired attributes other than improved serum half-life. For instance, the ability of the peptides to induce CTL activity can be enhanced by linkage to a sequence which contains at least one epitope that is capable of inducing a T helper cell response. Particularly preferred immunogenic peptides / T helper conjugates are linked by a spacer molecule. The spacer is typically comprised of relatively small, neutral molecules, such as amino acids or amino acid mimetics, which are substantiallyuncharged under physiological conditions. The spacers are typically selected from, e.g., Ala, Gly, or other neutral spacers of nonpolar amino acids or neutral polar amino acids. It will be understood that the optionally present spacer need not be comprised of the same residues and thus may be a hetero- or homo-oligomer. When present, the spacer will usually be at least one or two residues, more usually three to six residues. Alternatively, the peptide may be linked to the T helper peptide without a spacer.

[0182] The neoantigenic peptide may be linked to the T helper peptide either directly or via a spacer either at the amino or carboxy terminus of the peptide. The amino terminus of either the neoantigenic peptide or the T helper peptide may be acylated. Exemplary T helper peptides include tetanus toxoid 830-843, influenza 307-319, malaria circumsporozoite 382-398 and 378- 389.

[0183] Multiple neoantigenic peptides described herein can also be linked together, optionally by a spacer.Novel protein isoforms and antibodies binding thereof

[0184] The present disclosure provides novel protein isoforms that comprise non-exonic amino acid sequence, derived from variant transcripts that are identified as comprising non- exonic sequence and that are furthermore recurrent and translated. These variant transcripts are expressed and can generate protein isoforms that are stable in cells. The protein isoforms in some cases present modified functions compared to canonical isoforms.

[0185] Transcripts were assembled using RNA-Seq data in TCGA (the Cancer Genome Atlas) and CCLE (Broad Institute Cancer Cell Line Encyclopedia) (described in section EXAMPLES). Types of splicing alteration observed include exon skipping, intron retention and use of alternative splice donor or acceptor sites. In these variant transcripts, the non-exonic sequence can act as a donor (in 5’ position) or as an acceptor (in 3’ acceptor) and correspondingly the exon can be acceptor or donor. Non-exonic / exon splicing thus results in the incorporation of parts of the non-exonic sequence (i.e., “non-protein-coding” genome sequence) into the coding sequence of the variant transcript, thereby exposing non-exonic sequence (non-coding genomic sequences) to the translation machinery. When the non-exonic sequence is a TE, these variant transcripts can also be referred to as JET (Junction Exon TE).

[0186] The variant transcripts identified herein include an ORF (open reading frame), i.e. they are the part of a reading frame that has the ability to be translated into a polypeptide or protein. When the non-exonic sequence is acceptor, the ORF of the variant transcript is canonical (i.e. the same as the canonical transcript), whereas when the non-exonic sequence is the donor the ORF can be canonical (generally ORF1) or can be shifted by 1 or 2 nucleotides (generally ORFs 2 and 3 respectively) as compared to ORF1. The variant transcripts include not only the junction of non-exonic and exonic sequences (corresponding to the JET) but can also further include exon(s), upstream of the fusion breakpoint (between the exon and the non-exonic sequence) if the exon is donor or downstream of the fusion breakpoint if the non-exonic sequence is donor, corresponding to the various transcript isoforms. In any of the embodiments herein, the non- exonic sequence can be a TE, optionally intronic TE, optionally a short interspersed nuclear element (SINE), or a long interspersed nuclear element (LINE), or a long terminal repeat (LTRs). In some embodiments, the exonic sequence can be from a gene of phylostratum 1.

[0187] More particularly, the present disclosure provides novel protein isoforms of the amino acid sequences of Table 1 A, and variant transcripts of the nucleotide sequences of Table 1A which are referred to in Tables 1-4. Table 1A provides the canonical gene from which the variant transcript is derived. The variant transcript may have the same function as the canonical gene or in most cases a different function. Table IB provides detailed reference regarding transmembrane isoforms, which are expected to be displayed on the cell surface and thus are particularly good candidates for antibodies, TCRs or CARs are described herein. Some of these protein isoforms were also detected by mass-spectrometry proteomics. Table 3 provides a subset of protein isoforms that are significantly positively or negatively associated with cancer survival, and are particularly useful for diagnosing cancer or determining prognosis of cancer. Table 4 provides a subset of protein isoforms associated with cancer genes, e.g. derived from a gene annotated as cancer driver, tumor suppressor, or oncogene. 12 JETs in cancer driver genes are also correlated with cancer survival in at least 1 TCGA indication..

[0188] Protein isoforms, or fragments thereof, or neoantigenic peptides thereof, containing the non-exonic amino acid sequence as herein defined are ectopically expressed in a host cell, for example a tumor cell lines (such as Hela, CHO, etc.). In cases where the protein is transmembrane, the stability and proper integration within the plasma membrane can be confirmed. An expression vector can be designed containing a polynucleotide encoding the protein isoforms, or fragments thereof, or neoantigenic peptides thereof, containing non-exonic amino acid sequence, and a tag sequence (such as FLAG or HA) thus giving rise to a taggedfusion protein. The fusion-containing proteins or fragments or peptides also contain an epitopetag and can thus be detected by flow cytometry or microscopy using commercially available anti -tag antibodies. The sequences encoding the tag may be 5’ or 3’ to the protein coding sequence of the polynucleotide. In the case of transmembrane proteins, preferably the tag is located extracellularly.

[0189] Targeted sequencing experiments can also be performed to amplify and detect the variant transcripts in additional tumor specimens or cell lines. This can be performed through conventional PCR, quantitative real-time PCT or SMRT full-length transcripts sequencing (PACBIO technology).

[0190] In some embodiments, cell surface exposure of the non-exonic amino acid sequence can be predicted in silico based on the predicted topology of the protein isoform. Experimental validation of cell surface expression of the non-exonic amino acid sequence can also be performed.

[0191] In some embodiments, the transmembrane protein isoforms may be categorized as:(a) Integral membrane proteins(b) Type I single-pass proteins (positioned such that their carboxyl-terminus is towards the cytosol): selected transcripts are those derived from fusions in which TE acts as a donor (fusion TE->exon) and, in some cases, this fusion is preceded by a second fusion in which the TE is an acceptor (fusion exon-TE). In the later scenario the transcript is generated by a double-fusion “exon-TE-exon” or “metafusion” with a resulting transcript including a TE exonisated sequence flanked by two canonll exons.(c)Type II single-pass proteins (which have their amino-terminus towards the cytosol): selected transcripts are those derived from fusions in which TE acts as an acceptor (fusion exon->TE) and, in some cases, this fusion is followed by a second fusion in which the TE is a donor (fusion TE->exon). In the later scenario the transcript is generated by a double-fusion “exon-TE-exon” or “metafusion” with a resulting transcript including a TE exonisated sequence flanked by two canonical exons.(d) Multi-pass or poly- transmembrane proteins (the polypeptide chain crosses the membrane multiple times) : the TE sequence may act as donor or acceptor and / or may be part of a metafusion. The breakpoint (or one of the two breakpoints in the case of the metafusions) is located in the extracellular side of the membrane between two transmemble helices.(e) Integral monotopic proteins (integral membrane proteins that are attached to only one side of the membrane and do not span the whole way across): selected transcripts are those derived from fusions in which TE acts as an acceptor (fusion exon->TE), as a donor (fusion TE->exon) or as both (metafusions)(f) Peripheral membrane proteins (adhered or associated to the cell membrane): selected transcripts are those derived from fusions in which TE acts as an acceptor (fusion exon->TE), as a donor (fusion TE->exon) or as both (metafusions).

[0192] While the above-described examples of proteins are described with reference to TEs, the same applies to other non-exonic sequence (e.g., insertion of intron sequence or other noncoding sequence).

[0193] Typically, the novel protein isoform according to the present disclosure is expressed in more than 1 %, notably more than 5 %, and typically more than 10% of patient samples, e.g. tumor samples from cancer patients.

[0194] Typically, the protein isoform according to the present disclosure is expressed at higher levels in disease samples, e.g., tumor samples, as compared to normal samples. More particularly, the protein isoform is preferably expressed in less than 20%, notably less than 10 %, less than 5 % or less than 1 % of the normal samples.

[0195] The present disclosure also encompasses variants of novel protein isoforms having at least 50 %, notably at least 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, identity with any one of the amino acid sequences of Table 1A. Typically in said variants, the non-exonic amino acid sequence is preserved and has at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the non-exonic amino acid sequence of any of the amino acid sequences of Table 1A. Most preferably said variants protein isoforms proteins do not match any annotated polypeptide or protein in normal proteome databases such as UniProt.

[0196] The novel protein isoforms as herein disclosed can also be modified by extending or decreasing the compound's amino acid sequence, e.g., by the addition or deletion of amino acids. The protein isoforms can also be modified by altering the order or composition of certain residues, it being readily appreciated that certain amino acid residues essential for biological activity, e.g., those at critical contact sites or conserved residues, may generally not be altered without an adverse effect on biological activity. The non-critical amino acids need not be limited to those naturally occurring in proteins, such as L-a-amino acids, or their D-isomers,but may include non-natural amino acids as well, such as P-y-5-amino acids, as well as many derivatives of L-a-amino acids.

[0197] Typically, a series of peptides with single amino acid substitutions are employed to determine the effect of electrostatic charge, hydrophobicity, etc. on binding. The substitutions may be homo-oligomers or hetero-oligomers. The number and types of residues which are substituted or added depend on the spacing necessary between essential contact points and certain functional attributes which are sought (e.g., hydrophobicity versus hydrophilicity). In any event, such substitutions should employ amino acid residues or other molecular fragments chosen to avoid, for example, steric and charge interference which might disrupt binding.

[0198] Amino acid substitutions are typically of single residues. Substitutions, deletions, insertions or any combination thereof may be combined to arrive at a final peptide. Substitutional variants are those in which at least one residue of a peptide has been removed and a different residue inserted in its place. Such substitutions are generally made in accordance with the previously shown Table A when it is desired to finely modulate the characteristics of the peptide.

[0199] The present disclosure also encompasses antibodies or antigen-binding fragments thereof, e.g., antigen binding domains, that bind to a protein isoform as above defined or to a fragment thereof, notably to a neoantigenic tumor sequence thereof (or epitope) of a length at least 4, 5, 6 7, or 8 amino acids, with a dissociation constant (Kd) of about 2 x 10'7M or less. In certain embodiments, the Kd is about 2 x 10'7M or less, about 1 x 10'7M or less, about 9 x 10'8M or less, about 1 x 10'8M or less, about 9 x 10'9M or less, about 5 x 10'9M or less, about 4 x 10'9M or less, about 3 x 10'9or less, about 2 x 10'9M or less, or about 1 x 10'9M or less , or about 1 x IO'10M or less, about 1 x 10'11M or less, or about 1 x 10'12M or less. In certain non-limiting embodiments, the Kd is about 3 x 10'9M or less. In certain non-limiting embodiments, the Kd is from about 1 x 10'9M to about 3 x 10'7M. In certain non-limiting embodiments, the Kdis from about 1.5 x 10'9M to about 3 x 10'7M. In certain non-limiting embodiments, the Kd is from about 1.5 x 10'9M to about 2.7 x 10'7M. In certain non-limiting embodiments, the Kd is from about 1.5 x IO'10M to about 2.7 x 10'7M. In certain non-limiting embodiments, the Kd is from about 1.5 x 10'12M to about 2.7 x 10'7M.

[0200] Binding of the antigen-binding domain (for example, a Fv or an analog thereof) can be confirmed by, for example, enzyme-linked immunosorbent assay (ELISA), radioimmunoassay (RIA), FACS analysis, bioassay (e.g, growth inhibition), Western Blot assayor fluorescent microscopy. Each of these assays generally detect the presence of proteinantibody complexes of particular interest by employing a labeled reagent (e.g, an antibody, or an Fv) specific for the complex of interest. For example, the Fv can be radioactively labeled and used in a radioimmunoassay (RIA) (see, for example, Weintraub, B., Principles of Radioimmunoassays, Seventh Training Course on Radioligand Assay Techniques, The Endocrine Society, March, 1986, which is incorporated by reference herein). The radioactive isotope can be detected by such means as the use of a g counter or a scintillation counter or by autoradiography.

[0201] For microscopy, the labeled reagent (e.g, an antibody, or an Fv) can be either directly conjugated to a fluorophore or recognized by a fluorophore-conjugated secondary antibody directed against the labeled reagent. Non-limiting examples of fluorophores, also called fluorescent dyes, include derivatives of cyanine (e.g. Cy3) or rhodamine (e.g, TRITC) or fluorescein (e.g, FITC).

[0202] In certain embodiments, the extracellular antigen-binding domain is labeled with a fluorescent marker. Non-limiting examples of fluorescent markers include green fluorescent protein (GFP), blue fluorescent protein (e.g, EBFP, EBFP2, Azurite, and mKalamal), cyan fluorescent protein (e.g, ECFP, Cerulean, and CyPet), and yellow fluorescent protein (e.g, YFP, Citrine, Venus, and YPet).

[0203] In some embodiments, the antigen binding domain as herein disclosed binds to a fragment (or a neoantigenic peptide sequence) of the amino acid sequence (or an epitope) of a protein isoform as herein described, which comprises at least a non-exonic amino acid sequence. In some embodiments, the peptide sequence from the herein described protein isoform overlaps the breakpoint between, the non-exonic amino acid sequence and the exonic amino acid sequence. In other embodiments, the peptide sequence is derived from a pure non- exonic sequence, e.g. TE sequence. In yet other embodiments, the peptide sequence is encoded by a non-exonic amino acid sequence that is encoded by a non-canonical ORF downstream of the junction between the non-exonic sequence and the exonic sequence.

[0204] In some embodiments, the antigen binding domain according to the present disclosure binds a neoantigenic peptide sequence from any one of the novel protein isoforms as herein disclosed or fragment thereof, wherein said neoantigenic peptide sequence: is from any one of the amino acid sequences of Table 1 A or a fragment thereof and comprises at least a sequence derived from the non-exonic amino acid sequence,optionally (i) a fragment encoded by a part of the ORF that overlaps the breakpoint between, the non-exonic sequence and an exonic sequence or, optionally (ii) a pure non-exonic amino acid sequence, e.g., TE sequence; or optionally (iii) is non-exonic amino acid sequence encoded by a non-canonical ORF downstream of the junction between the non-exonic sequence and the exonic sequence.

[0205] In some embodiments, the peptide sequence is from an extracellular portion of the protein isoform.

[0206] In certain embodiments, the antigen-binding domain comprises an antigen binding portion of a TCR.

[0207] In certain embodiments, the antigen-binding domain comprises an antigen binding portion of an antibody or a fragment thereof. In certain embodiments, the antigen-binding domain comprises a heavy chain variable region (VH) and / or a light chain variable region (VL) of an antibody. In certain embodiments, the antigen-binding domain comprises a single-chain variable fragment (scFv).

[0208] In certain embodiments, the antigen-binding domain comprises a heavy chain-only antibodies (VHH).

[0209] In certain embodiments, the antigen-binding domain comprises a Fab, which is optionally crosslinked. In certain embodiments, the antigen-binding domain comprises a F(ab)2. In certain embodiments, any of the foregoing molecules can be comprised in a fusion protein with a heterologous sequence to form the antigen-binding domain.

[0210] In certain embodiments, the extracellular antigen-binding domain is derived from a scFv, Fab, or antibody of murine, human or camelid (e.g., lama) origin.

[0211] In some embodiments, the antigen-binding fragment comprises one, two or three CDRs of an antibody, e.g. of a VH or VL. In some embodiments, the antigen-binding fragments comprises four, five, or six CDRs of an antibody, e.g. VH and VL.Peptide production., polynucleotides and vectors

[0212] Proteins or peptides may be made by any technique known to those of skill in the art, including the expression of proteins, polypeptides or peptides through standard molecular biological techniques, the isolation of proteins or peptides from natural sources, or the chemical synthesis of proteins or peptides. The nucleotide and protein, polypeptide and peptide sequencescorresponding to various genes have been previously disclosed, and may be found at computerized databases known to those of ordinary skill in the art. One such database is the National Center for Biotechnology Information's Genbank and GenPept databases located at the National Institutes of Health website. The coding regions for known genes may be amplified and / or expressed using the techniques disclosed herein or as would be known to those of ordinary skill in the art. Alternatively, various commercial preparations of proteins, polypeptides and peptides are known to those of skill in the art.

[0213] In a further aspect the present disclosure provides a nucleic acid (e.g. polynucleotide) encoding a protein isoform, fragment thereof, or neoantigenic peptide thereof as described herein. The polynucleotide may be selected froINA, cDNA, PNA, CNA, RNA, either single- and / or double-stranded, or native or stabilized forms of polynucleotides, such as for example polynucleotides with a phosphorothiate backbone, or combinations thereof and it may or may not contain introns so long as it codes for the protein isoform, fragment thereof, or neoantigenic peptide. Only polypeptides or peptides that contain naturally occurring amino acid residues joined by naturally occurring peptide bonds are encodable by a polynucleotide.

[0214] A still further aspect of the disclosure provides an expression vector capable of expressing a protein isoform, fragment thereof, or neoantigenic peptide as herein disclosed. Expression vectors for different cell types are well known in the art and can be selected without undue experimentation. Generally, the DNA is inserted into an expression vector, such as a plasmid, in proper orientation and correct reading frame for expression. The expression vector will comprise the appropriate heterologous transcriptional and / or translational regulatory control nucleotide sequences recognized by the desired host. The polynucleotide encoding the protein isoform, fragment thereof, or neoantigenic peptide may be linked to such heterologous regulatory control nucleotide sequences or may be non-adjacent yet operably linked to such heterologous regulatory control nucleotide sequences. The vector is then introduced into the host cell through standard techniques. Guidance can be found for example in Sambrook et al. (1989) Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, N.Y.

[0215] Suitable host cells are known in the art. Methods of producing such protein isoform, fragment thereof, or neoantigenic peptides are also contemplated, involving culturing the host cells and isolating the desired protein isoform, fragment thereof, or neoantigenic peptide from the cell or culture medium.Antigen presenting cells (APCs)

[0216] The present disclosure also encompasses a population of antigen presenting cells that have been pulsed with one or more of the peptides as previously defined and / or comprise a polynucleotide encoding a protein isoform, fragment thereof, or neoantigenic peptide as described herein, and / or obtainable in a method as previously described. Preferably, the antigen presenting cells are dendritic cell (DCs) or artificial antigen presenting cells (aAPCs) (see Neal, Lillian R et al. “The Basics of Artificial Antigen Presenting Cells in T Cell-Based Cancer Immunotherapies.” Journal of immunology research and therapy vol. 2,1 (2017): 68-79). Dendritic cells (DC) are professional antigen-presenting cells (APC) that have an extraordinary city to stimulate naive T-cells and initiate primary immune responses to pathogens. Indeed, the main role of mature DCs are to sense antigens and produce mediators that activate other immune cells, particularly T cells. DCs are potent stimulators for lymphocyte activation as they express MHC molecules that trigger TCRs (signal 1) and co-stimulatory molecules (signal 2) on T cells. Additionally, DCs also secrete cytokines that support T cell expansion. T cells require presented antigen in the form of a processed peptide to recognize foreign pathogens or tumor. Presentation of peptide epitopes derived from pathogen / tumor proteins is achieved through MHC molecules. MHC class I (MHC -I) and MHC class II (MHC -II) molecules present processed peptides to CD8+ T cells and CD4+ T cells, respectively. Importantly, DCs home to inflammatory sites containing abundant T cell populations to foster an immune response. Thus, DCs can be a crucial component of any immunotherapeutic approach, as they are intimately involved with the activation of the adaptive immune response. In the context of vaccines, DC therapy can enhance T cell immune responses to a desired target in healthy volunteers or patients with infectious disease or cancer. In one embodiment, APCS are artificial APC, which are genetically modified to express the desired T-cell co-stimulatory molecules, human HLA alleles and / or cytokines. Such artificial antigen presenting cells (aAPC) are able to provide the requirements for adequate T-cell engagement, co-stimulation, as well as sustained release of cytokines that allow for controlled T-cell expansion. These cells are not subject to the constraints of time and limited availability and can be stored in small aliquots for subsequent use in generating T-cell lines from different donors, thus representing an off the shelf reagent for immunotherapy applications. Expression of potent co-stimulatory signals on these aAPC endows this system with higher efficiency lending to increased efficacy of adoptive immunotherapy. Furthermore, aAPC can be engineered to express genes directing release ofspecific cytokines to facilitate the preferential expansion of desirable T-cell subsets for adoptive transfer; such as long lived memory T-cells (see for review Hasan AH et al., Artificial Antigen Presenting Cells: An Off the Shelf Approach for Generation of Desirable T-Cell Populations for Broad Application of Adoptive Immunotherapy; Adv Genet Eng. 2015; 4(3): 130, Kim JV, Latouche JB, Riviere I, Sadelain M. The ABCs of artificial antigen presentation. Nat Biotechnol. 2004;22:403-410 or Wang C, Sun W, Ye Y, Bomba HN, Gu Z. Bioengineering of Artificial Antigen Presenting Cells and Lymphoid Organs. Theranostics 2017; 7(14):3504- 3516.).

[0217] Typically, the dendritic cells are autologous dendritic cells that are pulsed with a neoantigenic peptide as herein disclosed. The peptide may be any suitable peptide that gives rise to an appropriate T-cell response. The antigen-presenting cell (or stimulator cell) typically has an MHC class I or II molecule on its surface, and in one embodiment is substantially incapable of itself loading the MHC class I or II molecule with the selected antigen. The MHC class I or II molecule may readily be loaded with the selected antigen in vitro.

[0218] As an alternative the antigen presenting cell may comprise an expression construct encoding a neoantigenic peptide as herein disclosed. The polynucleotide may be any suitable polynucleotide as previously defined and it is preferred that it is capable of transducing the dendritic cell, thus resulting in the presentation of a peptide and induction of immunity.

[0219] Thus the present disclosure encompasses a population of APCs than can be pulsed or loaded with the neoantigenic peptide as herein disclosed, genetically modified (via DNA or RNA transfer) to express at least one neoantigenic peptide as herein disclosed, or that comprise an expression construct encoding a neoantigenic peptide of the present disclosure. Typically the population of APCs is pulsed or loaded, modified to express or comprises at least one, at least 5, at least 10, at least 15, or at least 20 different neoantigenic peptide or expression construct encoding it.

[0220] The present disclosure also encompasses compositions comprising APCs as herein disclosed. APCs can be suspended in any known physiologically compatible pharmaceutical carrier, such as cell culture medium, physiological saline, phosphate-buffered saline, cell culture medium, or the like, to form a physiologically acceptable, aqueous pharmaceutical composition. Parenteral vehicles include sodium’ chloride solution, Ringer's dextrose, dextrose and sodium chloride, lactated Ringer's. Other substances may be added as desired such as antimicrobials. As used herein, a “carrier” refers to any substance suitable as a vehicle fordelivering an APC to a suitable in vitro or in vivo site of action. As such, carriers can act as an excipient for formulation of a therapeutic or experimental reagent containing an APC. Preferred carriers are capable of maintaining an APC in a form that is capable of interacting with a T cell. Examples of such carriers include, but are not limited to water, phosphate buffered saline, saline, Ringer's solution, dextrose solution, serum-containing solutions, Hank's solution and other aqueous physiologically balanced solutions or cell culture medium. Aqueous carriers can also contain suitable auxiliary substances required to approximate the physiological conditions of the recipient, for example, enhancement of chemical stability and isotonicity. Suitable auxiliary substances include, for example, sodium acetate, sodium chloride, sodium lactate, potassium chloride, calcium chloride, sorbitan monolaurate, triethanolamine oleate, and other substances used to produce phosphate buffer, Tris buffer, and bicarbonate buffer.Vaccine Compositions

[0221] The present disclosure further encompasses a vaccine or immunogenic composition capable of raising a specific T-cell response comprising: one or more protein isoforms, fragment thereof, or neoantigenic peptides as herein described, one or more polynucleotides encoding a protein isoform, fragment thereof, or neoantigenic peptide as herein described; and / or

[0222] a population of antigen presenting cells (such as autologous dendritic cells or artificial APC) as described above.

[0223] In some embodiments, a protein isoform, fragment thereof, or neoantigenic peptide, or a polynucleotide encoding such a protein isoform, fragment thereof, or neoantigenic peptide, which is encoded by tumor specific variant transcripts, e.g. as disclosed in Tables 3 or 4 as are used in vaccine compositions as per the present disclosure. Said neoantigenic peptide can be referred to herein as tumor specific peptides. Preferably also polynucleotides encoding tumor specific peptides are used as per the present disclosure.

[0224] In other embodiments, a protein isoform, fragment thereof, or neoantigenic peptide, or a polynucleotide encoding such a protein isoform, fragment thereof, or neoantigenic peptide is useful in vaccine composittions for treating a disease associated with the gene function listed in any of Tables 1A, 3, or 4.

[0225] A vaccine or immunogenic composition can contain between 1 and 20 neoantigenic peptides, more preferably 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 different neoantigenic peptides, further preferred 6, 7, 8, 9, 10 11, 12, 13, or 14 different neoantigenic peptides, and most preferably 12, 13 or 14 different neoantigenic peptides.

[0226] The neoantigenic peptide(s) may be linked to a carrier protein. Where the composition contains two or more neoantigenic peptides, the two or more (e.g. 2-25) peptides may be linearly linked by a spacer molecule as described above, e.g. a spacer comprising 2-6 nonpolar or neutral amino acids.

[0227] In one embodiment of the present disclosure the different protein isoforms, fragments thereof, or neoantigenic peptides, encoding polynucleotides, vectors, or APCs are selected so that one vaccine or immunogenic composition comprises multiple multiple neoantigenic peptides capable of associating with different MHC molecules, such as different MHC class I molecules. Preferably, such neoantigenic peptides are capable of associating with the most frequently occurring MHC class I molecules, e.g. different fragments capable of associating with at least 2 preferred, more preferably at least 3 preferred, even more preferably at least 4 preferred MHC class I molecules. In some embodiments, the compositions comprise peptides, encoding polynucleotides, vectors, or APCs capable of associating with one or more MHC class II molecules. The MHC is optionally HLA -A, -B, -C, -DP, -DQ, or -DR.

[0228] The vaccine or immunogenic composition is capable of raising a specific cytotoxic T- cells response and / or a specific helper T-cell response.

[0229] A vaccine composition is to be understood as meaning a composition for generating immunity for the prophylaxis and / or treatment of diseases. Accordingly, vaccines are medicines which comprise or generate antigens and are intended to be used in humans or animals for generating specific defense and protective substance by vaccination. An “immunogenic composition” is to be understood as meaning a composition that comprises or generates antigen(s) and is capable of eliciting an antigen-specific humoral or cellular immune response, e.g. T-cell response.

[0230] In a preferred embodiment, the neoantigenic peptide according to the disclosure is 8 or 9 residues long, or from 13 to 25 residues long. When the peptide is less than 20 residues, in order to have a peptide better suited for in vivo immunization, said neoantigenic peptide, isoptionally flanked by additional amino acids to obtain an immunization peptide of more amino acids, usually more than 20.

[0231] Pharmaceutical compositions (i.e., the vaccine or immunogenic composition) comprising a a protein isoform, fragment thereof, or neoantigenic peptide as herein described may be administered to an individual already suffering from cancer. In therapeutic applications, compositions are administered to a patient in an amount sufficient to elicit an effective CTL response to the novel antigen and to cure or at least partially arrest symptoms and / or complications. An amount adequate to accomplish this is defined as "therapeutically effective dose." Amounts effective for this use will depend on, e.g., the peptide composition, the manner of administration, the stage and severity of the disease being treated, the weight and general state of health of the patient, and the judgment of the prescribing physician, but generally range for the initial immunization (that is for therapeutic or prophylactic administration) from about 1.0 pg to about 50,000 pg of peptide for a 70 kg patient, followed by boosting dosages or from about 1.0 pg to about 10,000 pg of peptide pursuant to a boosting regimen over weeks to months depending upon the patient's response and condition by measuring specific CTL activity in the patient's blood. It must be kept in mind that the peptide and compositions of the present invention may generally be employed in serious disease states, that is, life-threatening or potentially life threatening situations, especially when the cancer has metastasized. In such cases, in view of the minimization of extraneous substances and the relative nontoxic nature of the peptide, it is possible and may be felt desirable by the treating physician to administer substantial excesses of these peptide compositions.

[0232] For therapeutic use of cancer vaccines or immunogenic compositions, administration should begin at the detection or surgical removal of tumors. This is followed by boosting doses until at least symptoms are substantially abated and for a period thereafter.

[0233] The vaccine or immunogenic compositions for therapeutic treatment are intended for parenteral, topical, nasal, oral or local administration. Preferably, the pharmaceutical compositions are administered parenterally, e.g., intravenously, subcutaneously, intradermally, or intramuscularly. The compositions may be administered at the site of surgical excision to induce a local immune response to the tumor.

[0234] The vaccine or immunogenic composition may be a pharmaceutical composition which additionally comprises a pharmaceutically acceptable adjuvant, immunostimulatory agent, stabilizer, carrier, diluent, excipient and / or any other materials well known to thoseskilled in the art. Such materials should be non-toxic and should not interfere with the efficacy of the active ingredient. The carrier is preferably an aqueous carrier but its precise nature of the carrier or other material will depend on the route of administration. A variety of aqueous carriers may be used, e.g., water, buffered water, 0.9% saline, 0.3% glycine, hyaluronic acid and the like. These compositions may be sterilized by conventional, well known sterilization techniques, or may be sterile filtered. The resulting aqueous solutions may be packaged for use as is, or lyophilized, the lyophilized preparation being combined with a sterile solution prior to administration. The compositions may further contain pharmaceutically acceptable auxiliary substances as required to approximate physiological conditions, such as pH adjusting and buffering agents, tonicity adjusting agents, wetting agents and the like, for example, sodium acetate, sodium lactate, sodium chloride, potassium chloride, calcium chloride, sorbitan monolaurate, triethanolamine oleate, etc. See, for example, Butterfield, BMJ. 2015 22;350 for a discussion of cancer vaccines.

[0235] Example adjuvants that increase or expand the immune response of a host to an antigenic compound include emulsifiers, muramyl dipeptides, avridine, aqueous adjuvants such as aluminum hydroxide, chitosan-based adjuvants, saponins, oils, Amphigen, LPS, bacterial cell wall extracts, bacterial DNA, CpG sequences, synthetic oligonucleotides, cytokines and combinations thereof. Emulsifier include, for example, potassium, sodium and ammonium salts of lauric and oleic acid, calcium, magnesium and aluminum salts of fatty acids, organic sulfonates such as sodium lauryl sulfate, cetyltrhethyl ammonium bromide, glycerylesters, polyoxyethylene glycol esters and ethers, and sorbitan fatty acid esters and their polyoxyethylene, acacia, gelatin, lecithin and / or cholesterol. Adjuvants that comprise an oil component include mineral oil, a vegetable oil, or an animal oil. Othe’ adjuvants include Freund's Complet’ Adjuvant (FCA) or Freund's Incomplete Adjuvant (FIA). Cytokines useful as additional immunostimulatory agents include interferon alpha, interleukin-2 (IL-2), and granulocyte macrophage-colony stimulating factor (GM-CSF), or combinations thereof.

[0236] The concentration of peptides as herein described in the vaccine or immunogenic formulations can vary widely, i.e., from less than about 0.1 %, usually at or at least about 2 % to as much as 20 % to 50 % or more by weight, and will be selected primarily by fluid volumes, viscosities, etc., in accordance with the particular mode of administration selected.

[0237] The peptides as herein described may also be administered via liposomes, which target the peptides to a particular cells tissue, such as lymphoid tissue. Liposomes are also useful in increasing the half-life of the peptides. Liposomes include emulsions, foams, micelles,insoluble monolayers, liquid crystals, phospholipid dispersions, lamellar layers and the like. In these preparations the peptide to be delivered is incorporated as part of a liposome, alone or in conjunction with a molecule which binds to, e.g., a receptor prevalent among lymphoid cells, such as monoclonal antibodies which bind to the CD45 antigen, or with other therapeutic or immunogenic compositions. Thus, liposomes filled with a desired peptide of the invention can be directed to the site of lymphoid cells, where the liposomes then deliver the selected therapeutic / immunogenic peptide compositions. Liposomes for use in the invention are formed from standard vesicle-forming lipids, which generally include neutral and negatively charged phospholipids and a sterol, such as cholesterol. The selection of lipids is generally guided by consideration of, e.g., liposome size, acid lability and stability of the liposomes in the blood stream. A variety of methods are available for preparing liposomes, as described in, e.g., Szoka et al., Ann. Rev. Biophys. Bioeng. 9;467 (1980 ), USA U.S. Patent Nos. 4,235,871 , 4501728 USA 4,501,728, 4,837,028 , and 5,019,369 .

[0238] For targeting to the immune cells, a ligand to be incorporated into the liposome can include, e.g., antibodies or fragments thereof specific for cell surface determinants of the desired immune system cells. A liposome suspension containing a peptide may be administered intravenously, locally, topically, etc. in a dose which varies according to, inter alia, the manner of administration, the peptide being delivered, and the stage of the disease being treated.

[0239] For solid compositions, conventional or nanoparticle nontoxic solid carriers may be used which include, for example, pharmaceutical grades of mannitol, lactose, starch, magnesium stearate, sodium saccharin, talcum, cellulose, glucose, sucrose, magnesium carbonate, and the like. For oral administration, a pharmaceutically acceptable nontoxic composition is formed by incorporating any of the normally employed excipients, such as those carriers previously listed, and generally 10-95% of active ingredient, that is, one or more peptides of the invention, and more preferably at a concentration of 25%-75%.

[0240] For aerosol administration, the immunogenic peptides are preferably supplied in finely divided form along with a surfactant and propellant. Typical percentages of peptides are 0.01 %-20 % by weight, preferably l%-10%. The surfactant must, of course, be nontoxic, and preferably soluble in the propellant. Representative of such agents are the esters or partial esters of fatty acids containing from 6 to 22 carbon atoms, such as caproic, octanoic, lauric, palmitic, stearic, linoleic, linolenic, olesteric and oleic acids with an aliphatic polyhydric alcohol or its cyclic anhydride. Mixed esters, such as mixed or natural glycerides may be employed. The surfactant may constitute 0.1 %-20 % by weight of the composition, preferably 0.25-5 %. Thebalance of the composition is ordinarily propellant. A carrier can also be included as desired, as with, e.g., lecithin for intranasal delivery.

[0241] Cytotoxic T-cells (CTLs) recognize an antigen in the form of a peptide bound to an MHC molecule rather than the intact foreign antigen itself. The MHC molecule itself is located at the cell surface of an antigen presenting cell. Thus, an activation of CTLs is only possible if a trimeric complex of peptide antigen, MHC molecule, and antigen presenting cell (APC) is present. Correspondingly, it may enhance the immune response if not only the peptide is used for activation of CTLs, but if additionally APCs with the respective MHC molecule are added. Therefore, in some embodiments the vaccine or immunogenic composition according to the present disclosure alternatively or additionally contains at least one antigen presenting cell, preferably a population of APCs.

[0242] The vaccine or immunogenic composition may thus be delivered in the form of a cell, such as an antigen presenting cell, for example as a dendritic cell vaccine. The antigen presenting cells such as a dendritic cell may be pulsed or loaded with a protein isoform, fragment thereof, or neoantigenic peptide as herein disclosed, may comprise an expression construct encoding a a protein isoform, fragment thereof, or neoantigenic peptide as herein disclosed, or may be genetically modified (via DNA or RNA transfer) to express one, two or more of the herein disclosed neoantigenic peptides, for example at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 neoantigenic peptides.

[0243] Suitable vaccines or immunogenic compositions may also be in the form of DNA or RNA relating to a protein isoforms, fragments thereof, or neoantigenic peptides as described herein. For example, DNA or RNA encoding one or more protein isoforms, fragments thereof, or neoantigenic peptides may be used as the vaccine, for example by direct injection to a subject. For example, DNA or RNA encoding at least 2, 3, 4, 5, 6, 7, 8, 9 , 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 protein isoforms, fragments thereof, or neoantigenic peptides.

[0244] A number of methods are conveniently used to deliver the nucleic acids to the patient. For instance, the nucleic acid can be delivered directly, as "naked DNA". This approach is described, for instance, in Wolff et al., Science 247: 1465-1468 (1990 ) as well as USAU.S. Patent Nos. 5,580,859 and 5,589,466 . The nucleic acids can also be administered using ballistic delivery as described, for instance, in U.S. Patent No. 5,204,253. Particles comprised solely of DNA can be administered. Alternatively, DNA can be adhered to particles, such as gold particles.

[0245] The nucleic acids can also be delivered complexed to cationic compounds, such as cationic lipids. Lipid-mediated gene delivery methods are described, for instance, in WO 96 / 18372; WO 93 / 24640; Mannino & Gould-Fogerite, BioTechniques 6(7): 682-691 (1988); U.S. Pat No. 5,279,833; WO 91 / 06309; and Feigner et al., Proc. Natl. Acad. Sci. USA 84: 7413- 7414 (1987 ).

[0246] Delivery systems may optionally include cell-penetrating peptides, nanoparticulate encapsulation, virus like particles, liposomes, or any combination thereof. Cell penetrating peptides include TAT peptide, herpes simplex virus VP22, transportan, Antp. Liposomes may be used as a delivery system. Listeria vaccines or electroporation may also be used.

[0247] The protein isoform, fragment thereof, or neoantigenic peptide may also be delivered via a bacterial or viral vector containing DNA or RNA sequences which encode one or more protein isoforms, fragments thereof, or neoantigenic peptides. The DNA or RNA may be delivered as a vector itself or within attenuated bacteria virus or live attenuated virus, such as vaccinia or fowlpox. This approach involves the use of vaccinia virus as a vector to express nucleotide sequences that encode the peptide of the invention. Upon introduction into an acutely or chronically infected host or into a noninfected host, the recombinant vaccinia virus expresses the immunogenic peptide, and thereby elicits a host CTL response. Vaccinia vectors and methods useful in immunization protocols are described in, e.g., U.S. Patent No. 4,722,848. Another vector is BCG (Bacille Calmette Guerin). BCG vectors are described in Stover et al. (Nature 351 :456-460 (1991 )). A wide variety of other vectors useful for therapeutic administration or immunization of the peptides of the invention, e.g., Salmonella typhivectors and the like, will be apparent to those skilled in the art from the description herein.

[0248] An appropriate means of administering nucleic acids encoding the neoantigenic peptides as herein described involves the use of minigene constructs encoding multiple epitopes. To create a DNA sequence encoding the selected CTL epitopes (minigene) for expression in human cells, the amino acid sequences of the epitopes are reverse translated. A human codon usage table is used to guide the codon choice for each amino acid. These epitopeencoding DNA sequences are directly adjoined, creating a continuous polypeptide sequence. To optimize expression and / or immunogenicity, additional elements can be incorporated into the minigene design. Examples of amino acid sequence that could be reverse translated and included in the minigene sequence include: helper T lymphocyte, epitopes, a leader (signal) sequence, and an endoplasmic reticulum retention signal. In addition, MHC presentation ofCTL epitopes may be improved by including synthetic (e.g. poly-alanine) or naturally- occurring flanking sequences adjacent to the CTL epitopes.

[0249] The minigene sequence is converted to DNA by assembling oligonucleotides that encode the plus and minus strands of the minigene. Overlapping oligonucleotides (30-100 bases long) are synthesized, phosphorylated, purified and annealed under appropriate conditions using well known techniques. The ends of the oligonucleotides are joined using T4 DNA ligase. This synthetic minigene, encoding the CTL epitope polypeptide, can then cloned into a desired expression vector.

[0250] Standard regulatory sequences well known to those of skill in the art are included in the vector to ensure expression in the target cells. Thus, the DNA or RNA encoding the protein isoform, fragment thereof, or neoantigenic peptide(s) may typically be operably linked to one or more of:

[0251] Promoter that can be used to drive nucleic acid molecule expression, e.g. constitutive or ubiquitous promoters such as CMV (notably human cytomegalovirus immediate early promoter (hCMV-IE)), CAG, CBh, PGK, SV40, RSV, Ferritin heavy or light chains, etc. For brain expression, the following promoters can be used: Synapsinl for all neurons, CaMKIIalpha for excitatory neurons, GAD67 or GAD65 or VGAT for GABAergic neurons, etc. Promoters used to drive RNA synthesis can include: Pol III promoters such as U6 or HI . The use of a Pol II promoter and intronic cassettes can be used to express guide RNA (gRNA). Typically, the promoter includes a down-stream cloning site for minigene insertion. For examples of suitable promoters sequences, see notably U.S. Patent Nos. 5,580,859 and 5,589,466.

[0252] Transcriptional transactivators or other enhancer elements, which can also increase transcription activity, e.g. the regulatory R region from the 5' long terminal repeat (LTR) of human T-cell leukemia virus type 1 (HTLV-1) (which when combined with a CMV promoter has been shown to induce higher cellular immune response).

[0253] Translation optimizing sequences e.g. a Kozak sequence flanking the AUG initiator codon (ACCAUGG) within mRNA, and codon optimization.

[0254] Additional vector modifications may be desired to optimize minigene expression and immunogenicity. In some cases, introns are required for efficient gene expression, and one or more synthetic or naturally-occurring introns could be incorporated into the transcribed region of the minigene. The inclusion of mRNA stabilization sequences can also be considered for increasing minigene expression. It has recently been proposed that immunostimulatorysequences (ISSs or CpGs) play a role in the immunogenicity of DNA vaccines. These sequences could be included in the vector, outside the minigene coding sequence, if found to enhance immunogenicity.

[0255] In some embodiments, a bicistronic expression vector, to allow production of the minigene-encoded epitopes and a second protein included to enhance or decrease immunogenicity can be used.

[0256] DNA vaccines or immunogenic compositions as herein described can be enhanced by co-delivering cytokines that promote cell-mediated immune responses, such as IL-2, IL- 12, IL- 18, GM-CSF and IFNy. CXC chemokines such as IL-8, and CC chemokines such as macrophage inflammatory protein (MlP)-la, MIP-3a, MIP-3P, and RANTES, may increase the potency of the immune response. DNA vaccine immunogenicity can also be enhanced by codelivering plasmid-encoded cytokine-inducing molecules (e.g. LelF), co-stimulatory and adhesion molecules, e.g. B7-1 (CD80) and / or B7-2 (CD86). Helper (HTL) epitopes could be joined to intracellular targeting signals and expressed separately from the CTL epitopes. This would allow direction of the HTL epitopes to a cell compartment different than the CTL epitopes. If required, this could facilitate more efficient entry of HTL epitopes into the MHC class II pathway, thereby improving CTL induction. In contrast to CTL induction, specifically decreasing the immune response by co-expression of immunosuppressive molecules (e.g. TGF- P) may be beneficial in certain diseases.

[0257] Once an expression vector is selected, the minigene is cloned into the polylinker region downstream of the promoter. This plasmid is transformed into an appropriate E. coli strain, and DNA is prepared using standard techniques. The orientation and DNA sequence of the minigene, as well as all other elements included in the vector, are confirmed using restriction mapping and DNA sequence analysis. Bacterial cells harboring the correct plasmid can be stored as a master cell bank and a working cell bank.

[0258] Purified plasmid DNA can be prepared for injection using a variety of formulations. The simplest of these is reconstitution of lyophilized DNA in sterile phosphate-buffer saline (PBS). A variety of methods have been described, and new techniques may become available. As noted above, nucleic acids are conveniently formulated with cationic lipids. In addition, glycolipids, fusogenic liposomes, peptides and compounds referred to collectively as protective, interactive, non-condensing (PINC) could also be complexed to purified plasmidDNA to influence variables such as stability, intramuscular dispersion, or trafficking to specific organs or cell types.

[0259] Vaccines or immunogenic compositions comprising peptides may be administered in combination with vaccines or immunogenic compositions comprising polynucleotide encoding the peptides. For example, administration of peptide vaccine and DNA vaccine may be alternated in a prime-boost protocol. For example, priming with a peptide immunogenic composition and boosting with a DNA immunogenic composition is contemplated, as is priming with a DNA immunogenic composition and boosting with a peptide immunogenic composition.

[0260] The present disclosure also encompasses a method for producing a vaccine composition comprising the steps of:

[0261] Optionally, identifying at least one neoantigenic peptide according to the method as previously described;

[0262] producing said at least one neoantigenic peptide, at least one polypeptide encoding neoantigenic peptide(s), or at least a vector comprising said polypeptide(s) as described herein; and

[0263] optionally adding physiologically acceptable buffer, excipient and / or adjuvant and producing a vaccine with said at least one neoantigenic peptide, polypeptide or vector.

[0264] Another aspect of the present disclosure, is a method for producing a DC vaccine, wherein said DCs present at least one neoantigenic peptide as herein disclosed.Antibodies, TCRs, CARs and derivatives thereof

[0265] The present disclosure also relates to an antibody, including an antigen-binding fragment or an antigen-binding domain, that specifically binds a novel protein isoform as described herein, and most preferably a protein isoform of any of the amino acid sequences of Table 1 A, or a neoantigenic peptide typically in association with an MHC or HLA molecule, as herein disclosed.

[0266] In some embodiments, said antibody or antigen-binding fragment thereof comprises or consists of an antigen-binding domain (that bind a protein isoform) as described herein.

[0267] Typically, said antibody, or antigen-binding fragment thereof binds a protein isoform, fragment thereof, or neoantigenic peptide, said peptide typically in association with an MHC or HLA molecule or a protein isoform, as previously defined, with a dissociation constant (Kd) of about 2 x 10'7M or less. In certain embodiments, the Kd is about 2 x 10'7M or less, about 1 x 10'7M or less, about 9 x 10'8M or less, about 1 x 10'8M or less, about 9 x 10'9M or less, about 5 x 10’9M or less, about 4 x 10'9M or less, about 3 x 10'9or less, about 2 x 10'9M or less, or about 1 x 10'9M or less. In certain non-limiting embodiments, the Kd is about 3 x 10"9M or less. In certain non-limiting embodiments, the Kd is from about 1 x 10'9M to about 3 x 10'7M. In certain non-limiting embodiments, the Kd is from about 1.5 x 10'9M to about 3 x 10'7M. In certain non-limiting embodiments, the Kd is from about 1.5 x 10'9M to about 2.7 x IO’7M.

[0268] To promote the infiltration and recognition of tumor cells by lymphocytes T (LT), another strategy consists in using antibodies capable of recognizing more than one antigenic target simultaneously and more particularly two antigenic targets simultaneously. There are many formats of bispecific antibodies. BiTE (bi-specific T-cell engager) are the first to have been developed. These are proteins of fusion consisting of two scFvs (variable domains heavy VH and light VL chains) from two antibodies linked by a binding peptide: one recognizes the LT marker (CD3+) and the other a tumor antigen. The goal is to favor recruitment and activation of LTs in contact with tumor, thus leading to cell lysis tumor (See for review Patrick A. Baeuerle and Carsten Reinhardt; Bispecific T-Cell Engaging Antibodies for Cancer Therapy; Cancer Res 2009; 69: (12). June 15, 2009 ; and Galaine et al., Innovations & Therapeutiques en Oncologie, vol. 3-n°3-7, mai-aout 2017).

[0269] In a particular embodiment, said antibody is thus a bi-specific T-cell engager that targets a protein isoform as herein defined, and in particular that comprises an antigen binding domain as pre“iously d”fined.

[0270] The term "antibody" herein is used in the broadest sense and includes polyclonal and monoclonal antibodies, including intact antibodies and functional (antigen-binding) antibody fragments, including fragment antigen binding (Fab) fragments, F(ab')2 fragments, Fab' fragments, Fv fragments, recombinant IgG (rlgG) fragments, variable heavy chain (VH) regions capable of specifically binding the antigen, single chain antibody fragments, including single chain variable fragments (scFv), and single domain antibodies (e.g., VHH antibodies, sdAb, sdFv, nanobody) fragments. The term encompasses genetically engineered and / or otherwise variants modified forms of immunoglobulins, such as intrabodies, peptibodies, chimericantibodies, fully human antibodies, humanized antibodies, and heteroconjugate antibodies, multispecific, e.g., bispecific, antibodies, diabodies, triabodies, and tetrabodies, tandem di- scFv, tandem tri-scFv. Unless otherwise stated, the term "antibody" should be understood to encompass functional antibody and fragments thereof. The term also encompasses intact or full- length antibodies, including antibodies of any class or sub-class, including IgG and sub-classes thereof, IgGl, IgG2, IgG3, IgG4, IgM, IgE, IgA, and IgD. In some embodiments, the antibody comprises a light chain variable domain and a heavy chain variable domain, e.g. in an scFv format.

[0271] Antibodies include variant polypeptide species that have one or more amino acid substitutions, insertions, or deletions in the native amino acid sequence, provided that the antibody retains or substantially retains its specific binding function. Conservative substitutions of amino acids are well known and described above.

[0272] The present disclosure further includes a method of producing an antibody, or antigenbinding fragment thereof, comprising a step of selecting antibodies that bind to a protein isoform, fragment thereof, or neoantigenic peptide as herein defined, typically in association with an MHC or HLA molecule, or that bind a protein isoform, fragment thereof, or neoantigenic peptide as herein defined with a dissociation constant (Kd) of about 2 x 10'7M or less. In certain embodiments, the Kd is about 2 x 10'7M or less, about 1 x 10'7M or less, about 9 x 10'8M or less, about 1 x 10'8M or less, about 9 x 10'9M or less, about 5 x 10'9M or less, about 4 x 10'9M or less, about 3 x 10'9or less, about 2 x 10'9M or less, or about 1 x 10'9M or less , or about 1 x IO'10M or less, or about 1 x 10'11M or less, or about 1 x 10'12M or less.

[0273] In certain embodiments, the antibody is of murine, human or camelid (e.g., lama) origin.

[0274] In some embodiments, the antibodies are selected from a library of human antibody sequences and the library is contacted with a protein isoform, fragment thereof, or neoantigenic peptide described herein. In some embodiments, the antibodies are generated by immunizing an animal with a protein isoform of any one of the amino acid sequences of Table 1A, as previously defined, or a fragment thereof (in particular with the extracellular portion), or a neoantigenic peptide thereof, followed by the selection step.

[0275] Antibodies including chimeric, humanized or human antibodies can be further affinity matured and selected as described above. Humanized antibodies contain rodent-sequence derived CDR regions; typically the rodent CDRs are engrafted into a human framework, andsome of the human framework residues may be back-mutated to the original rodent framework residue to preserve affinity, and / or one or a few of the CDR residues may be mutated to increase affinity. Fully human antibodies have no murine sequence, and are typically produced via phage display technologies of human antibody libraries, or immunization of transgenic mice whose native immunoglobin loci have been replaced with segments of human immunoglobulin loci.

[0276] Antibodies produced by said method, as well as immune cells expressing such antibodies or fragments thereof are also encompassed by the present disclosure.

[0277] The present disclosure also encompasses pharmaceutical compositions comprising one or more antibodies, including antigen-binding fragments, as herein disclosed alone or in combination with at least one other agent, such as a stabilizing compound, which may be administered in any sterile, biocompatible pharmaceutical carrier and optionally formulated with formulated with sterile pharmaceutically acceptable buffer(s), diluent(s), and / or excipient(s).

[0278] The present disclosure also encompasses a recombinant T cell receptor (TCR) that targets a neoantigenic protein isoform, fragment thereof, or neoantigenic peptide as herein defined in association with an MHC or HLA molecule.

[0279] The present disclosure further includes a method of producing a TCR, or an antigenbinding fragment thereof, comprising a step of selecting TCRs that bind to a protein isoform, fragment thereof, or neoantigenic peptide as herein defined, optionally in association with an MHC or HLA molecule, with a dissociation constant (Kd) of about 2 x 10'7M or less. In certain embodiments, the Kd is about 2 x 10'7M or less, about 1 x 10'7M or less, about 9 x 10'8M or less, about 1 x 10'8M or less, about 9 x 10'9M or less, about 5 x 10'9M or less, about 4 x 10"9M or less, about 3 x 10'9or less, about 2 x 10'9M or less, or about 1 x 10'9M or less, or about 1 x 10'10M or less, or about 1 x 10'11M or less or about 1 x 10'12M or less.

[0280] Nucleic acid encoding the TCR can be obtained from a variety of sources, such as by polymerase chain reaction (PCR) amplification of naturally occurring TCR DNA sequences, followed by expression of antibody variable regions, followed by the selecting step described above. In some embodiments, the TCR is obtained from T-cells isolated from a patient, or from cultured T-cell hybridomas. In some embodiments, the TCR clone for a target antigen has been generated in transgenic mice engineered with human immune system genes (e.g., the human leukocyte antigen system, or HLA). See, e.g., tumor antigens (see, e.g., Parkhurst et al. (2009) Clin Cancer Res. 15: 169-180 and Cohen et al. (2005) J Immunol. 175:5799-5808. In someembodiments, phage display is used to isolate TCRs against a target antigen (see, e.g., Varela- Rohena et al. (2008) Nat Med. 14: 1390-1395 and Li (2005) Nat Biotechnol. 23:349-354.

[0281] The present disclosure also provides an ex vivo method for producing a T cell comprising a TCR that specifically binds any of the novel protein isoforms, or fragments thereof or neoantigenic peptides thereof, comprising: (a) stimulating T cells from a blood sample obtained from a patient with a composition comprising said protein isoform, fragment or peptide, or polynucleotide encoding said protein isoform, fragment or peptide, to prime, activate, and expand T-cells, optionally CD4+ T-cells, CD8+ T cells, or effector or central memory T cells, (b) optionally prior to step (a), the method comprises obtaining a sample of cells or tissue from the patient, optionally by leukapheresis or tumor biopsy, and confirming expression of said protein isoform or fragment thereof in said sample, (c) optionally subsequent to step (a), the method comprises confirming the specificity and functionality, optionally cytokine production activity and cytolytic activity, of the induced T-cells. Such T-cells are useful in cell therapy, e.g., cell therapy for treating cancer.

[0282] A "T cell receptor" or "TCR" refers to a molecule that contains a variable a and P chains (also known as TCRa and TCRb, respectively) or a variable y and 5 chains (also known as TCRg and TCRd, respectively) and that is capable of specifically binding to an antigen peptide bound to a MHC receptor. In some embodiments, the TCR is in the aP form. Typically, TCRs that exist in aP and y5 forms are generally structurally similar, but T cells expressing them may have distinct anatomical locations or functions. A TCR can be found on the surface of a cell or in soluble form. Generally, a TCR is found on the surface of T cells (or T lymphocytes) where it is generally responsible for recognizing antigens bound to major histocompatibility complex (MHC) molecules through its extracellular binding domain. In some embodiments, a TCR also can contain a constant domain, a transmembrane domain and / or a short cytoplasmic tail (see, e.g., Janeway et ah, Immunobiology: The Immune System in Health and Disease, 3 rd Ed., Current Biology Publications, p. 4:33, 1997). For example, in some aspects, each chain of the TCR can possess one N-terminal immunoglobulin variable domain, one immunoglobulin constant domain, a transmembrane region, and a short cytoplasmic tail at the C-terminal end. In some embodiments, a TCR is associated with invariant proteins of the CD3 complex involved in mediating signal transduction. Unless otherwise stated, the term "TCR" should be understood to encompass functional TCR fragments thereof. The term also encompasses intact or full-length TCRs, including TCRs in the aP form or y5 form.

[0283] Thus, for purposes herein, reference to a TCR includes any TCR or functional fragment, such as an antigen-binding portion of a TCR that binds to a specific antigenic peptide bound in an MHC molecule, i.e. MHC -peptide complex. An "antigen-binding portion" or” antigen-binding fragment" of a TCR, which can be used interchangeably, refers to a molecule that contains a portion of the structural domains of a TCR, but that binds the antigen (e.g. MHC- peptide complex) to which the full TCR binds. In some cases, an antigen-binding portion contains the variable domains of a TCR, such as variable a chain and variable P chain of a TCR, sufficient to form a binding site for binding to a specific MHC -peptide complex, such as generally where each chain contains three complementarity determining regions.

[0284] In some embodiments, the variable domains of the TCR chains associate to form loops, or complementarity determining regions (CDRs) analogous to immunoglobulins, which confer antigen recognition and determine peptide specificity by forming the binding site of the TCR molecule and determine peptide specificity. Typically, like immunoglobulins, the CDRs are separated by framework regions (FRs) {see, e.g., lores et al., Pwc. Nat'lAcad. Sci. U.S.A. 87:9138, 1990; Chothia et al., EMBO J. 7:3745, 1988; see also Lefranc et al., Dev. Comp. Immunol. 27:55, 2003). In some embodiments, CDR3 is the main CDR responsible for recognizing processed antigen, although CDR1 of the alpha chain has also been shown to interact with the N-terminal part of the antigenic peptide, whereas CDR1 of the beta chain interacts with the C-terminal part of the peptide. CDR2 is thought to recognize the MHC molecule. In some embodiments, the variable region of the P-chain can contain a further hypervariability (HV4) region.

[0285] In some embodiments, the TCR chains contain a constant domain. For example, like immunoglobulins, the extracellular portion of TCR chains (e.g., a-chain, P-chain) can contain two immunoglobulin domains, a variable domain (e.g., Va or Vp; typically amino acids 1 to 116 based on Kaba“ numbering Kabat et al., "Sequences of Proteins of Immunological Interest, US Dept. Health and Human Services, Public Health Service National Institutes of Health, 1991, 5th ed.) at the N-terminus, and one constant domain (e.g., a-chain constant domain or Ca, typically amino acids 117 to 259 based on Kabat, P-chain constant domain or Cp, typically amino acids 117 to 295 based on Kabat) adjacent to the cell membrane. For example, in some cases, the extracellular portion of the TCR formed by the two chains contains two membrane- proximal constant domains, and two membrane-distal variable domains containing CDRs. The constant domain of the TCR domain contains short connecting sequences in which a cysteine residue forms a disulfide bond, making a link between the two chains. In some embodiments,a TCR may have an additional cysteine residue in each of the a and P chains such that the TCR contains two disulfide bonds in the constant domains.

[0286] In some embodiments, the TCR chains can contain a transmembrane domain. In some embodiments, the transmembrane domain is positively charged. In some cases, the TCR chains contain a cytoplasmic tail. In some cases, the structure allows the TCR to associate with other molecules like CD3. For example, a TCR containing constant domains with a transmembrane region can anchor the protein in the cell membrane and associate with invariant subunits of the CD3 signaling apparatus or complex.

[0287] Generally, CD3 is a multi-protein complex that can possess three distinct chains (y, 5, and a) in mammals and the ^-chain. For example, in mammals the complex can contain a CD3y chain, a CD35 chain, two CD3s chains, and a homodimer of CD3(^ chains. The CD3y, CD35, and CD3s chains are highly related cell surface proteins of the immunoglobulin superfamily containing a single immunoglobulin domain. The transmembrane regions of the CD3y, CD35, and CD3s chains are negatively charged, which is a characteristic that allows these chains to associate with the positively charged T cell receptor chains. The intracellular tails of the CD3y, CD35, and CD3s chains each contain a single conserved motif known as an immunoreceptor tyrosine -based activation motif or ITAM, whereas each CD3^ chain has three. Generally, ITAMs are involved in the signaling capacity of the TCR complex. These accessory molecules have negatively charged transmembrane regions and play a role in propagating the signal from the TCR into the cell. The CD3- and (^-chains, together with the TCR, form what is known as the T cell receptor complex.

[0288] In some embodiments, the TCR may be a heterodimer of two chains a and P (or optionally y and 5) or it may be a single chain TCR construct. In some embodiments, the TCR is a heterodimer containing two separate chains (a and P chains or y and 5 chains) that are linked, such as by a disulfide bond or disulfide bonds.

[0289] While T-cell receptors (TCRs) are transmembrane proteins and do not naturally exist in soluble form, antibodies can be secreted as well as membrane bound. Importantly, TCRs have the advantage over antibodies that they in principle can recognize peptides generated from all degraded cellular proteins, both intra- and extracellular, when presented in the context of MHC molecules. Thus TCRs have important therapeutic potential.

[0290] The present disclosure also relates to soluble T-cell receptors (sTCRs) that contain the antigen recognition part directed against a protein isoform, fragment thereof, or neoantigenicpeptide thereof as herein disclosed (see notably Walseng E, Walchli S, Fallang L-E, Yang W, Vefferstad A, Areffard A, et al. (2015) Soluble T-Cell Receptors Produced in Human Cells for Targeted Delivery. PLoS ONE 10(4): eOl 19559). In a particular embodiment, the soluble TCR can be fused to an antibody fragment directed to a T cell antigen, optionally wherein the targeted antigen is CD3 or CD16 (see for example Boudousquie, Caroline et al. “Polyfunctional response by ImmTAC (IMCgplOO) redirected CD8+ and CD4+ T cells.” Immunology vol. 152,3 (2017): 425-438. doi: 10.1111 / imm.l2779).

[0291] In certain embodiments, the present disclosure encompasses Recombinant HLA- independent (or non-HLA restricted) T cell receptors (referred to as“HI-TCRs”) that bind to a protein isoform, fragment thereof, or neoantigenic peptide thereof as described herein (including a fragment of any one of the amino acid sequences of Table 1A) in an HLA- independent manner. “HI-TCRs” as herein intended and which are well-suited to the present invention are described in International Application No. WO 2019 / 157454. Thus, typically HI- TCRs according to the present disclosure comprise an antigen binding chain that comprises: (a) an antigen-binding domain (as previously defined) that binds to an antigen in an HLA- independent manner, for example, an antigen-binding fragment of an immunoglobulin variable region; and (b) a constant domain that is capable of associating with (and consequently activating) a CD3(^ polypeptide. Because typically TCRs bind antigen in a HLA-dependent manner, the antigen-binding domain that binds in an HLA-independent manner is heterologous. Preferably, the antigen-binding domain or fragment thereof comprises: (i) an antigen-binding domain comprising or consisting of an heavy chain variable region (VH) of an antibody and / or (ii) a light chain variable region (VL) of an antibody. The constant domain of the TCR comprises, for example, a native or modified TRAC polypeptide, or a native or modified TRBC polypeptide. The constant domain of the TCR comprises, for example, at least one native TCR constant domain (e.g., alpha or beta) or fragment thereof. Unlike chimeric antigen receptors, which typically themselves comprise an intracellular signaling domain, the HI-TCR does not directly produce an activating signal; instead, the antigen-binding chain associates with and consequently activates a CD3(^ polypeptide. The immune cells comprising the recombinant TCR provide superior activity when the antigen has a low density on the cell surface of less than about 10,000 molecules per cell, or less than about 5,000 molecules per cell.

[0292] The CD3^ polypeptide is, for example, a native CD3(^ polypeptide or a modified CD3^ polypeptide. The CD3(^ polypeptide is optionally fused to an intracellular domain of a costimulatory molecule or a fragment thereof. Alternatively, the antigen binding domainoptionally comprises a co-stimulatory region, e.g. intracellular domain, that is capable of stimulating an immunoresponsive cell upon the binding of the antigen binding chain to the antigen. Example co-stimulatory molecules include CD28, 4-1BB, 0X40, ICOS, DAP-10, fragments thereof, or a combination thereof.

[0293] In some embodiments, the recombinant HI-TCR is expressed by a transgene that is integrated at an endogenous gene locus of the immunoresponsive cell, for example, a CD35 locus, a CD3s locus, a CD247 locus, a B2M locus, a TRAC locus, a TRBC locus, a TRDC locus and / or a TRGC locus. In most embodiments, expression of the recombinant HI-TCR is driven from the endogenous TRAC or TRBC gene locus. In some embodiments, the transgene encoding a portion of the recombinant HI-TCR is integrated into the endogenous TRAC and / or TRBC locus in a manner that disrupts or abolishes the endogenous expression of a TCR comprising a native TCR a chain and / or a native TCR P chain. This disruption prevents or eliminates mispairing between the recombinant TCR and a native TCR a chain and / or a native TCR P chain in the immunoresponsive cell. The endogenous gene locus may also comprise a modified transcription terminator region, for example, a TK transcription terminator, a GCSF transcription terminator, a TCRA transcription terminator, an HBB transcription terminator, a bovine growth hormone transcription terminator, an SV40 transcription terminator, and a P2A element.

[0294] In some embodiments of the present disclosure, the recombinant TCR and typically the HI-TCR comprises an extracellular antigen-binding domain which is capable of dimerizing with a second extracellular antigen-binding domain. Typically, the second extracellular antigen-binding domain binds a tumor antigen, preferably wherein the tumor antigen is selected from pHER95, CD19, MUC16, MUC1, CAIX, CEA, CD8, CD7, CD10, CD20, CD22, CD30, CD70, CLL1, CD33, CD34, CD38, CD41, CD44, CD49f, CD56, CD74, CD133, CD138, EGP- 2, EGP-40, EpCAM, Erb-B2, Erb-B3, Erb-B4, FBP, Fetal acetylcholine receptor, folate receptor-a, GD2, GD3, HER-2, hTERT, IL-13R-a2, K-light chain, KDR, LeY, LI cell adhesion molecule, MAGE-A1, Mesothelin, MAGEA3, p53, MARTI, GP100, Proteinase3 (PR1), Tyrosinase, Survivin, hTERT, EphA2, NKG2D ligands, NY-ESO-1, oncofetal antigen (h5T4), PSCA, PSMA, R0R1, TAG-72, VEGF-R2, WT-1, BCMA, CD123, CD44V6, NKCS1, EGF1R, EGFR-VIII, CD99, CD70, ADGRE2, CCR1, LILRB2, LILRB4, PRAME, and ERBB.

[0295] The present disclosure also encompasses a chimeric antigen receptor (CAR) which is directed against a novel protein isoform as herein disclosed and in particular a protein isoform of any of the amino acid sequences of Table 1A. In preferred embodiments, the CAR comprisesan antigen-binding domain as previously defined. CARs are fusion proteins comprising an antigen-binding domain, typically derived from an antibody, linked to the signalling domain of the TCR complex. CARs can be used to direct immune cells, such as T-cells or NK T cells, against a protein isoform, fragment thereof, or neoantigenic peptide as previously defined with a suitable antigen-binding domain selected.

[0296] The antigen-binding domain of a CAR is typically based on a scFv (single chain variable fragment) derived from an antibody. In addition to an N-terminal, extracellular antibody-binding domain, CARs typically may comprise a hinge domain, which functions as a spacer to extend the antigen-binding domain away from the plasma membrane of the immune effector cell on which it is elssed, a transmembrane (TM) domain, an intracellular signalling domain (e.g. the signalling domain from the zeta chain of the CD3 molecule (CD3Q of the TCR complex, or an equivalent) and optionally one or more co- stimulatory domains which may assist in signalling or functionality of the cell expressing the CAR. Signalling domains from co-stimulatory molecules including CD28, OX-40 (CD134), and 4-1BB (CD137) can be added alone (second generation) or in combination (third generation) to enhance survival and increase proliferation of CAR modified T cells. Potential co-stimulatory domains also include ICOS-1, CD27, GITR, and DAP 10.

[0297] Thus, the CAR may include:

[0298] In its extracellular portion, one or more antigen binding molecules, such as one or more antigen-binding fragment, domain, or portion of an antibody, or one or more antibody variable domains, and / or antibody molecules, and typically one or more antigen-binding domain as previously defined.

[0299] In its transmembrane portion, a transmembrane domain derived from human T cell receptor-alpha or -beta chain, a CD3 zeta chain, CD28, CD3-epsilon, CD45, CD4, CD5, CD8, CD9, CD16, CD22, CD33, CD37, CD64, CD80, CD86, CD134, CD137, ICOS, CD154, or a GITR. In some embodiments, the transmembrane domain is derived from CD28, CD8 or CD3- zeta.

[0300] One or more co-stimulatory domains, such as co-stimulatory domains derived from human CD28, 4-1BB (CD137), ICOS-1, CD27, OX 40 (CD137), DAP10, and GITR (AITR). In some embodiments, the CAR comprises co-stimulating domains of both CD28 and 4-1BB.

[0301] In its intracellular signalling domain, an intracellular signalling domain comprising one or more IT AMs, for example, the intracellular signalling domain is CD3-zeta, or a variantthereof lacking one or two ITAMs (e.g. ITAM3 and ITAM2), or the intracellular signalling domain is derived from FcsRIy.

[0302] The CAR can be designed to recognize protein isoform, fragment thereof, or neoantigenic peptide alone or in association with an HLA or MHC molecule.

[0303] The moi eties used to bind to antigen include three general categories, either singlechain antibody fragments (scFvs) derived from antibodies, Fab’s selected from libraries, or natural ligands that engage their cognate receptor (for the first-generation CARs). Successful examples in each of these categories are notably reported in Sadelain M, Brentjens R, Riviere I. The basic principles of chimeric antigen receptor (CAR) design. Cancer discovery. 2013; 3(4):388-398 (see notably table 1) and are included in the present application.

[0304] Antibodies include chimeric, humanized or human antibodies, and can be further affinity matured and selected as described above. Chimeric or humanized scFv’s derived from rodent immunoglobulins (e.g. mice, rat) are commonly used, as they are easily derived from well-characterized monoclonal antibodies. Humanized antibodies contain rodent-sequence derived CDR regions; typically the rodent CDRs are engrafted into a human framework, and some of the human framework residues may be back-mutated to the original rodent framework residue to preserve affinity, and / or one or a few of the CDR residues may be mutated to increase affinity. Fully human antibodies have no murine sequence, and are typically produced via phage display technologies of human antibody libraries, or immunization of transgenic mice whose native immunoglobin loci have been replaced with segments of human immunoglobulin loci. Variants of the antibodies can be produced that have one or more amino acid substitutions, insertions, or deletions in the native amino acid sequence, wherein the antibody retains or substantially retains its specific binding function. Conservative substitutions of amino acids are well known and described above. Further variants may also be produced that have improved affinity for the antigen.

[0305] Typically, the CAR includes an antigen-binding domain as previously defined from an antibody molecule, such as a single-chain antibody fragment (scFv) derived from the variable heavy (VH) and variable light (VL) chains of a monoclonal antibody (mAb).

[0306] In some aspects, the antigen- binding, domain of the CAR is linked to one or more transmembrane and intracellular signaling domains. In some embodiments, the CAR includes a transmembrane domain fused to the extracellular domain of the CAR. In one embodiment, the transmembrane domain that is naturally associated with one of the domains in the CAR isused. In some instances, the transmembrane domain is selected or modified by amino acid substitution to avoid binding of such domains to the transmembrane domains of the same or different surface membrane proteins to minimize interactions with other members of the receptor complex.

[0307] The transmembrane domain in some embodiments is derived either from a natural or from a synthetic source. Where the source is natural, the domain can be derived from any membrane-bound or transmembrane protein. Transmembrane regions include those derived from (i.e. comprise at least the transmembrane region(s) of) the alpha, beta or zeta chain of the T-cell receptor, CD28, CD3 epsilon, CD45, CD4, CD5, CD8, CD9, CD16, CD22, CD33, CD37, CD64, CD80, CD86, CD 134, CD137, CD154, ICOS or a GITR). The transmembrane domain can also be synthetic. In some embodiments, the transmembrane domain is derived from CD28, CD 8 or CD3-zeta.

[0308] In some embodiments, a short oligo- or polypeptide linker, for example, a linker of between 2 and 10 amino acids in length, is present and forms a linkage between the transmembrane domain and the cytoplasmic signaling domain of the CAR.

[0309] The CAR generally includes at least one intracellular signaling component or components. First generation CARs typically had the intracellular domain from the CD3 C,- chain, which is the primary transmitter of signals from endogenous TCRs. Second generation CARs typically further comprise intracellular signaling domains from various costimulatory protein receptors (e.g., CD28, 41BB (CD28), ICOS) to the cytoplasmic tail of the CAR to provide additional signals to the T cell. Co-stimulatory domains include domains derived from human CD28, 4-1BB (CD137), ICOS-1, CD27, OX 40 (CD137), DAP10, and GITR (AITR). Combinations of two co-stimulatory domains are contemplated, e.g. CD28 and 4-1BB, or CD28 and 0X40. Third generation CARs combine multiple signaling domains, such as CD3z-CD28- 4-1BB or CD3z-CD28-OX40, to augment potency.

[0310] The intracellular signaling domain can be from an intracellular component of the TCR complex, such as a TCR CD3+ chain that mediates T-cell activation and cytotoxicity, e.g., the CD3 zeta chain. Alternative intracellular signaling domains include FcsRIy. The intracellular signaling domain may comprise a modified CD3 zeta polypeptide lacking one or two of its three immunoreceptor tyrosine-based activation motifs (ITAMs), wherein the ITAMs are ITAM1, ITAM2 and ITAM3 (numbered from the N-terminus to the C-terminus). The intracellular signaling region of CD3-zeta is residues 22-164, of which IT AMI is located around amino acidresidues 61-89, ITAM2 around amino acid residues 100-128, and ITAM3 around residues 131- 159. Thus, the modified CD3 zeta polypeptide may have any one of IT AMI, ITAM2, or ITAM3 inactivated. Alternatively, the modified CD3 zeta polypeptide may have any two ITAMs inactivated, e.g. ITAM2 and ITAM3, or ITAM1 and ITAM2. Preferably, ITAM3 is inactivated, e.g. deleted. More preferably, ITAM2 and ITAM3 are inactivated, e.g. deleted, leaving ITAM1. For example, one modified CD3 zeta polypeptide retains only ITAM1 and the remaining CD3(^ domain is deleted (residues 90-164). As another example, ITAM1 is substituted with the amino acid sequence of ITAM3, and the remaining CD3^ domain is deleted (residues 90-164). See, for example, Bridgeman et al., Clin. Exp. Immunol. 175(2): 258-67 (2014); Zhao et al., J. Immunol. 183(9): 5563-74 (2009); Maus et al., WO 2018 / 132506; Sadelain et al., WO / 2019 / 133969, Feucht et al., Nat Med. 25(l):82-88 (2019).

[0311] Thus, in some aspects, the antigen binding domain is linked to one or more cell signaling modules. In some embodiments, cell signaling modules include CD3 transmembrane domain, CD3 intracellular signaling domains, and / or other CD transmembrane domains. The CAR can also further include a portion of one or more additional molecules such as Fc receptor y, CD8, CD4, CD25, or CD 16.

[0312] In some embodiments, upon ligation of the CAR, the cytoplasmic domain or intracellular signaling domain of the CAR activates at least one of the normal effector functions or responses of the corresponding non-engineered immune cell (typically a T cell). For example, the CAR can induce a function of a T cell such as cytolytic activity or T-helper activity, secretion of cytokines or other factors.

[0313] In some embodiments, the intracellular signaling domain(s) include the cytoplasmic sequences of the T cell receptor (TCR), and in some aspects also those of co-receptors that in the natural context act in concert with such receptor to initiate signal transduction following antigen-specific receptor engagement, and / or a variant of such molecules, and / or any synthetic sequence that has the same functional capability.

[0314] T cell activation is in some aspects described as being mediated by two classes of cytoplasmic signaling sequences: those that initiate antigen- dependent primary activation through the TCR (primary cytoplasmic signaling sequences), and those that act in an antigenindependent manner to provide a secondary or co- stimulatory signal (secondary cytoplasmic signaling sequences). In some aspects, the CAR includes one or both of such signaling components.

[0315] In some aspects, the CAR includes a primary cytoplasmic signaling sequence that regulates primary activation of the TCR complex either in a stimulatory way, or in an inhibitory way. Primary cytoplasmic signaling sequences that act in a stimulatory manner may contain signaling motifs which are known as immunoreceptor tyrosine -based activation motifs or ITAMs. Examples of IT AM containing primary cytoplasmic signaling sequences include those derived from TCR zeta, FcR gamma, FcR beta, CD3 gamma, CD3 delta, CD3 epsilon, CDS, CD22, CD79a, CD79b, and CD66d. In some embodiments, cytoplasmic signaling molecule(s) in the CAR contain(s) a cytoplasmic signaling domain, portion thereof, or sequence derived from CD3 zeta.

[0316] The CAR can also include a signaling domain and / or transmembrane portion of a costimulatory receptor, such as CD28, 4-1BB, 0X40, DAP10, and ICOS. In some aspects, the same CAR includes both the activating and costimulatory components; alternatively, the activating domain is provided by one CAR whereas the costimulatory component is provided by another CAR recognizing another antigen.

[0317] The CAR or other antigen-specific receptor can also be an inhibitory CAR (e.g. iCAR) and includes intracellular components that dampen or suppress a response, such as an immune response. Examples of such intracellular signaling components are those found on immune checkpoint molecules, including PD-1, CTLA4, LAG3, BTLA, 0X2R, TIM-3, TIGIT, LAIR- 1, PGE2 receptors, EP2 / 4 Adenosine receptors including A2AR. In some aspects, the engineered cell includes an inhibitory CAR including a signaling domain of or derived from such an inhibitory molecule, such that it serves to dampen the response of the cell. Such CARs are used, for example, to reduce the likelihood of off-target effects when the antigen recognized by the activating receptor, e.g, CAR, is also expressed, or may also be expressed, on the surface of normal cells.

[0318] Exemplary antigen receptors, including CARs and recombinant TCRs, as well as methods for engineering and introducing the receptors into cells, include those described, for example, in international patent application publication numbers W0200014257, WO2013126726, WO2012 / 129514, WO2014031687, WO2013 / 166321, WO2013 / 071154, W02013 / 123061, WO2019157454, U.S. patent application publication numbers US2002131960, US2013287748, US20130149337, U.S. Patent Nos.: 6,451,995, 7,446,190, 8,252,592, , 8,339,645, 8,398,282, 7,446,179, 6,410,319, 7,070,995, 7,265,209, 7,354,762, 7,446,191, 8,324,353, and 8,479,118, and European patent application number EP2537416, and / or those described by Sadelain et al., Cancer Discov. 2013 April; 3(4): 388-398; Davila etal. (2013) PLoS ONE 8(4): e61338; Turtle et al., Curr. Opin. Immunol., 2012 October; 24(5): 633-39; Wu et al., Cancer, 2012 March 18(2): 160-75. In some aspects, the genetically engineered antigen receptors include a CAR as described in U.S. Patent No.: 7,446,190, and those described in International Patent Application Publication No.: WO / 2014055668 Al.

[0319] The present disclosure also encompasses polynucleotides encoding antibodies, antigen-binding fragments or derivatives thereof, TCRs and CARs as previously described as well as vector comprising said polynucleotide(s).Immune cells

[0320] The present disclosure further encompasses an immune cell, notably an isolated immune cell which target one or more protein isoform, fragment thereof, or neoantigenic peptides as previously described. In more specific embodiments the present disclosure encompasses an immune cell, notably an isolated immune cell expressing a recombinant CAR or TCR as previously defined.

[0321] As used herein, the term “immune cell” includes cells that are of hematopoietic origin and that play a role in the immune response. Immune cells include lymphocytes, such as B cells and T cells, natural killer cells, myeloid cells, such as monocytes, macrophages, eosinophils, mast cells, basophils, and granulocytes.

[0322] As used herein, the term “T cell” includes cells bearing a T cell receptor (TCR), in particular TCR directed against a protein isoform, fragment thereof, or neoantigenic peptide as herein disclosed. T-cells according to the present disclosure can be selected from the group consisting of inflammatory T-lymphocytes, cytotoxic T-lymphocytes, regulatory T- lymphocytes, Mucosal-Associated Invariant T cells (MAIT), Y5 T cell, tumour infiltrating lymphocyte (TILs) or helper T- lymphocytes included both type 1 and 2 helper T cells and Thl7 helper cells. In another embodiment, said cell can be derived from the group consisting of CD4+ T- lymphocytes and CD8+ T-lymphocytes. Said immune cells may originate from a healthy donor or from a subject suffering from a cancer. In some embodiments, the immune cell is an allogenic or autologous cell. In some embodiments, the immune cell is selected from T cells, Natural Killer T cells, CD4+ / CD8+ T cells, TILs / tumor derived CD8 T cells, central memory CD8+ T cells, Treg, MAIT, Y5 T cells, human embryonic stem cells, and pluripotent stem cells from which lymphoid cells may be differentiated.

[0323] Immune cells can be extracted from blood or derived from stem cells. The stem cells can be adult stem cells, embryonic stem cells, more particularly non-human stem cells, cord blood stem cells, progenitor cells, bone marrow stem cells, induced pluripotent stem cells, totipotent stem cells or hematopoietic stem cells. Adipose cells can also be induced to form stem cells. Representative human cells are CD34+ cells.

[0324] T-cells can be obtained from a number of non-limiting sources, including peripheral blood mononuclear cells, bone marrow, lymph node tissue, cord blood, thymus tissue, tissue from a site of infection, ascites, pleural effusion, spleen tissue, and tumors. In certain embodiments, T-cells can be obtained from a unit of blood collected from a subject using any number of techniques known to the skilled person, such as FICOLL™ separation. In one embodiment, cells from the circulating blood of a subject are obtained by apheresis. In certain embodiments, T-cells are isolated from PBMCs. PBMCs may be isolated from buffy coats obtained by density gradient centrifugation of whole blood, for instance centrifugation through a LYMPHOPREP™ gradient, a PERCOLL™ gradient or a FICOLL™ gradient. T-cells may be isolated from PBMCs by depletion of the monocytes, for instance by using CD 14 DYNABEADS®. In some embodiments, red blood cells may be lysed prior to the density gradient centrifugation.

[0325] In another embodiment, said cell can be derived from a healthy donor, from a subject diagnosed with cancer. The cell can be autologous or allogeneic.

[0326] In allogeneic immune cell therapy, immune cells are collected from healthy donors, rather than the patient. Typically these are HLA matched to reduce the likelihood of graft vs. host disease. Alternatively, universal ‘off the shelf products that may not require HLA matching comprise modifications designed to reduce graft vs. host disease, such as disruption or removal of the TCRaP receptor. See Graham et al., Cells. 2018 Oct; 7(10): 155 for a review. Because a single gene encodes the alpha chain (TRAC) rather than the two genes encoding the beta chain, the TRAC locus is a typical target for removing or disrupting TCRaP receptor expression. Alternatively, inhibitors of TCRaP signalling may be expressed, e.g. truncated forms of CD3(^ can act as a TCR inhibitory molecule. Disruption or removal of HLA class I molecules has also been employed. For example, Torikai et al., Blood. 2013;122: 1341-1349 used ZFNs to knock out the HLA-A locus, while Ren et al., Clin. Cancer Res. 2017;23:2255- 2266 knocked out Beta-2 microglobulin (B2M), which is required for HLA class I expression. Ren et al. simultaneously knocked out TCRaP, B2M and the immune-checkpoint PD1. Generally, the immune cells are activated and expanded to be utilized in the adoptive celltherapy. The immune cells as herein disclosed can be expanded in vivo or ex vivo. The immune cells, in particular T-cells can be activated and expanded generally using methods known in the art. Generally the T-cells are expanded by contact with a surface having attached thereto an agent that stimulates a CD3 / TCR complex associated signal and a ligand that stimulates a costimulatory molecule on the surface of the T cells.

[0327] In one embodiment of the present disclosure, the immune cell can be modified to be directed to a protein isoform, fragment thereof, or neoantigenic peptide as previously defined. In a particular embodiment, said immune cell may express a recombinant antigen receptor directed to said protein isoform, fragment thereof, or neoantigenic peptide on its cell surface. By "recombinant" is meant an antigen receptor which is not encoded by the cell in its native state, i.e. it is heterologous, non-endogenous. Expression of the recombinant antigen receptor can thus be seen to introduce new antigen specificity to the immume cell, causing the cell to recognise and bind a previously described protein isoform, fragment thereof, or neoantigenic peptide. The antigen receptor may be isolated from any useful source. In some embodiments, the cells comprise one or more nucleic acids introduced via genetic engineering that encode one or more antigen receptors, wherein the antigen include at least one protein isoform, fragment thereof, or neoantigenic peptide as per the present disclosure.

[0328] Among the antigen receptors as per the present disclosure are genetically engineered T cell receptors (TCRs) and components thereof, as well as functional non-TCR antigen receptors, such as chimeric antigen receptors (CAR) as previously described.

[0329] Methods by which immune cells can be genetically modified to express a recombinant antigen receptor are well known in the art. A nucleic acid molecule encoding the antigen receptor may be introduced into the cell in the form of e.g. a vector, or any other suitable nucleic acid construct. Vectors, and their required components, are well known in the art. Nucleic acid molecules encoding antigen receptors can be generated using any method known in the art, e.g. molecular cloning using PCR. Antigen receptor sequences can be modified using commonly- used methods, such as site-directed mutagenesis.

[0330] In some embodiments of the present disclosure, the immune cell is a cell wherein (a) the SUV39H1 gene is inactivated, (b) the antigen-specific receptor is a modified TCR comprising a heterologous (or recombinant) antigen-binding domain as previously defined and a native TCR constant domain or fragment thereof, and the antigen-specific receptor is capable of activating a CD3 zeta polypeptide. For example, the immune cell may further comprise atleast one chimeric costimulatory receptor (CCR) and / or at least one chimeric antigen receptor, for example as previously defined.

[0331] In a related aspect, the immune cells, particularly if allogeneic, may be designed to reduce graft vs. host disease, such that the cells comprise inactivated (e.g. disrupted or deleted) TCRaP receptor. In such cases, the nucleic acid encoding the antigen-binding domain of the HI-TCR (typically as previously defined) is conveniently inserted into the endogenous TRAC locus and / or TRBC locus of the immune cell. The insertion of the HI-TCR nucleic acid sequence, or another smaller mutation, can disrupt or abolish the endogenous expression of a TCR comprising a native TCR alpha chain and / or a native TCR beta chain. The insertion or mutation may reduce endogenous TCR expression by at least about 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99%. Because a single gene encodes the alpha chain (TRAC) rather than the two genes encoding the beta chain, the TRAC locus is a typical target for reducing TCRaP receptor expression. Thus, the nucleic acid encoding the antigen-specific receptor (e.g. CAR or TCR) may be integrated into the TRAC locus at a location, preferably in the 5’ region of the first exon, that significantly reduces expression of a functional TCR alpha chain. See, e.g., Jantz et al., WO 2017 / 062451; Sadelain et al., WO 2017 / 180989; Torikai et al,. Blood, 119(2): 5697-705 (2012); Eyquem et al., Nature. 2017 Mar 2;543(7643): 113-117. Expression of the endogenous TCR alpha may be reduced by at least about 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99%. In such embodiments, expression of the nucleic acid encoding the antigen-specific receptor is optionally under control of the endogenous TCR- alpha or endogenous TCR-beta promoter.

[0332] Optionally, the immune cell also comprises a modified CD3 with a single active IT AM domain, and optionally the CD3 may further comprise one or more or two or more costimulatory domains. In some embodiments, the CD3 comprises two costimulatory domains, optionally CD28 and 4-1BB. The modified CD3 with a single active ITAM domain can comprise, for example, a modified CD3zeta intracellular signaling domain in which ITAM2 and ITAM3 have been inactivated, or ITAM1 and ITAM2 have been inactivated. In some embodiments, a modified CD3 zeta polypeptide retains only ITAM1 and the remaining CD3(^ domain is deleted (residues 90-164). As another example, ITAM1 is substituted with the amino acid sequence of ITAM3, and the remaining CD3(^ domain is deleted (residues 90-164).

[0333] The modified immune cells disclosed herein may comprise combinations of two or more, or three or more, or four or more, of the foregoing aspects.

[0334] For example, the modified immune cell is an immune cell wherein (a) the antigenspecific receptor is a modified TCR comprising a heterologous (or recombinant) antigenbinding domain (typically as previously defined) and a native TCR constant domain or fragment thereof, and the antigen-specific receptor is capable of activating a CD3 zeta polypeptide, and / or the antigen-specific receptor is a CAR, and optionally (b) the SUV39H1 gene is inactivated, and optionally (c) the immune cell comprises a modified CD3 with a single active ITAM domain, e.g. in which ITAM2 and ITAM3 have been inactivated, and optionally (d) the TCR is under control of an endogenous TRAC and / or IC promoter, and optionally (e) expression of native TCR-alpha chain and / or native TCR-beta chain are disrupted or abolished. In further embodiments, the cell may comprise at least one chimeric costimulatory receptor (CCR).

[0335] The present disclosure also relates to a method for providing an immune cell, and in particular a T cell population which targets a protein isoform, fragment thereof, or neoantigenic peptide as herein disclosed, in particular an immune cell and notably a T cell population expressing a TCR, notably a HLA Independent TCR (HI TCR) or a CAR as previously defined.

[0336] The T cell population may comprise CD8+ T cells, CD4+ T cells or CD8+ and CD4+ T cells.

[0337] Immune cell populations produced in accordance with the present disclosure may be enriched with immune cells that are specific to, i.e. target, the protein isoforms, fragments thereof, or neoantigenic peptides of the present disclosure and in particular the protein isoforms of any of the amino acid sequences of Table 1 A. That is, the immune cell population that is produced in accordance with the present disclosure will have an increased number of immune cells that target one or more protein isoform, fragment thereof, or neoantigenic peptide (i.e. enriched in clonotypes targeting the neoantigenic peptide). For example, the immune cell population of the disclosure will have an increased number of immune cells that target a protein isoform, fragment thereof, or neoantigenic peptide compared with the immune cells in the sample isolated from the subject. That is to say, the composition of the immune cell population will differ from that of a "native" immune cell population (i.e. a population that has not undergone the identification and expansion steps discussed herein), in that the percentage or proportion of immune cells that target a protein isoform, fragment thereof, or a neoantigenic peptide described herein will be increased.

[0338] The immune cell population according to the present disclosure may have at least about 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 1 1 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 or 100% T cells that target a protein isoform, fragment thereof, or neoantigenic peptide as herein disclosed. For example, the immune cell population may have about 0.2%-5%, 5%-10%, 10-20%, 20-30%, 30-40%, 40-50 %, 50-70% or 70-100% immune cells that target a protein isoform, fragment thereof, or neoantigenic peptide of the present disclosure.

[0339] Ari eta et al, “201 BNT221, an autologous neoantigen-specific T-cell product for adoptive cell therapy of metastatic ovarian cancer,” J. ImmunoTher. Cancer 2021;9:doi: 10.1136 / jitc-2021-SITC2021.201 describes an ex vivo stimulation protocol using synthetic neoantigenic peptides or mRNA encoding neoantigenic peptides to prime, activate, and expand T cells, including responsive CD4+ T cells, CD8+ T cells, effector memory phenotype cells, and central memory phenotype cells.

[0340] An expanded population of neoantigenic peptide / or protein isoform -reactive immune cells may have a higher activity than a population of immune cells not expanded, for example, when exposing those cells to a protein isoform, fragment thereof, or neoantigenic peptide thereof. Reference to "activity" may represent the response of the immune cell population to restimulation with a neoantigenic peptide (e.g. a peptide corresponding to the peptide used for expansion) or a mix of neoantigenic peptide or with a protein isoform as herein defined (or fragment thereof, typically extracellular fragment thereof) or with a mix of protein isoform (or fragment thereof, typically extracellular fragment thereof) . Suitable methods for assaying the response are known in the art. For example, cytokine production may be measured (e.g. IL2 or IFNy production may be measured). The reference to a "higher activity" includes, for example, a 1-5, 5-10, 10-20, 20-50, 50-100, 100-500, 500-1000-fold increase in activity. In one aspect the activity may be more than 1000-fold higher.

[0341] In a preferred embodiment present disclosure provides a plurality or population, i.e. more than one, of immune cells wherein the plurality of immune cells comprises a immune cell, notably a T cell, which recognizes a clonal neoantigenic peptide and a T cell which recognizes a different clonal neoantigenic peptide. As such, the present disclosure provides a plurality of immune cells, notably T cells, which recognize different clonal neoantigenic peptides. Different immune cells, notably T cells, in the plurality or population may alternatively have different TCRs which recognize different epitopes of the same neoantigenic peptide or protein isoform.

[0342] In a preferred embodiment the number of clonal neoantigenic peptides or protein isoforms or epitopes of one or more protein isoform(s) recognized by the plurality of T cells is from 2 to 1000. For example, the number of clonal neo-antigens recognized may be 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950 or 1000, preferably 2 to 100. There may be a plurality of immune cells, notably T cells, with different TCRs but which recognize the same clonal neo-antigen.

[0343] The immune cell and in particular the T cell population may be all or primarily composed of CD8+ T cells, or all or primarily composed of a mixture of CD8+ T cells and CD4+ T cells or all or primarily composed of CD4+ T cells.

[0344] In particular embodiments, the T cell population is generated from T cells isolated from patient with a disease associated with a gene function in Table 1 A, 3, or 4, e.g. a cancer patient, or a healthy donor. For example, the T cell population may be generated from T cells in a sample isolated from a tumor-bearing patient. The sample may be a tumor sample, a peripheral blood sample or a sample from other tissues of the subject.

[0345] In a particular embodiment the immune cell population is generated from a sample from the tissue in which the protein isoform or neoantigenic peptide is identified. In other words, the immune cell and notably the T cell population is isolated from a biological specimen derived from the tumor of a cancer patient. Such T cells are referred to herein as 'tumor infiltrating lymphocytes' (TILs).

[0346] T cells may be isolated using methods which are well known in the art. For example, T cells may be purified from single cell suspensions generated from samples on the basis of expression of CD3, CD4 or CD8. T cells may be enriched from samples by passage through a Ficoll-paque gradient.Oligonucleotide therapeutic agents

[0347] Oligonucleotide therapeutic agents are known in the art and include, for example, antisense oligonucleotide constructs, siRNAs, shRNAs, micro RNA (miRNA), saRNA, aptamers, ribozymes, and splice-switching antisense oligonucleotides (SSOs).

[0348] The term “antisense oligonucleotide” or “antisense nucleic acid” as used herein refers to a single-stranded oligonucleotide having a nucleobase sequence that is complementary to a corresponding segment of a target nucleic acid, e.g., a target genomic DNA sequence, pre- mRNA, or mRNA molecule.

[0349] Anti-sense oligonucleotides, including anti-sense RNA molecules and anti-sense DNA molecules, would act to directly block the translation of a gene transcript, for example, any of the variant transcripts disclosed herein, thus preventing protein translation or increasing mRNA degradation, in turn decreasing the level of the encoded protein isoform and thus its activity in a cell. For example, antisense oligonucleotides of at least about 15 bases and complementary to unique regions of the variant transcript can be synthesized, e.g., by conventional phosphodiester techniques and administered by e.g., intravenous injection or infusion. Methods for using antisense techniques for specifically inhibiting gene expression of genes whose sequence is known are well known in the art (see for example U.S. Pat. Nos. 6,566,135; 6,566, 131; 6,365,354; 6,410,323; 6,107,091; 6,046,321; and 5,981,732).

[0350] Small inhibitory RNAs (siRNAs) can also function as inhibitors of expression for use in the present disclosure. The expression of any gene, e.g., any of the variant transcripts described herein can be reduced by contacting a subject or cell with a small double stranded RNA (dsRNA), or a vector or construct causing the production of a small double stranded RNA, such that variant transcript expression is specifically inhibited (i.e. RNA interference or RNAi). Methods for selecting an appropriate dsRNA or dsRNA-encoding vector are well known in the art for genes whose sequence is known (see for example Tuschl, T. et al. (1999); Elbashir, S. M. et al. (2001); Hannon, GJ. (2002); McManus, MT. et al. (2002); Brummelkamp, TR. et al. (2002); U.S. Pat. Nos. 6,573,099 and 6,506,559; and International Patent Publication Nos. WO 01 / 36646, WO 99 / 32619, and WO 01 / 68836). All or parts of the phosphodiester bonds of the siRNAs of the disclosure are advantageously protected. This protection is generally implemented via the chemical route using methods that are known in the art. The phosphodiester bonds can be protected, for example, by a thiol or amine functional group or by a phenyl group. The 5'- and / or 3'- ends of the siRNAs of the disclosure are also advantageously protected, for example, using the technique described above for protecting the phosphodiester bonds. The siRNA sequences advantageously comprise at least twelve contiguous dinucleotides or their derivatives.

[0351] As used herein, the term "siRNA derivatives" with respect to the present nucleic acid sequences refers to any nucleic acid having a percentage of identity of at least 90% with erythropoietin or fragment thereof, preferably of at least 95%, as an example of at least 98%, and more preferably of at least 98%.

[0352] The siRNA can also be linked together with a loop to form shRNAs (short hairpin RNA) which can also function as inhibitors of expression for use in the present disclosure.

[0353] MicroRNAs (miRNAs) are small (about 21-23 nucleotides) noncoding RNAs that post transcriptionally regulating target gene expression through base pairing to partially complementary sites to prevent protein accumulation by repressing translation or by inducing mRNA degradation. These characteristics make them a possible tool for inhibiting protein translation. See, e.g., Juanjuan Zhao et al., “MicroRNA-7: a promising new target in cancer therapy” Cancer Cell International 2015; 15: 103. miRNA-122 is expressed in hepatocytes and interacts with hepatitis C virus (HCV) leading to proliferation of HCV. Luna et al., Cell 2015; 160: 1099-110.

[0354] Small activating RNAs (saRNAs) can also induce gene expression. These small dsRNA, typically about 21 nucleotides in length, target specific gene promoters and induce transcriptional gene activation. Li et al. (2006) . Proc Natl Acad Sci USA 103, 17337-17342 designed a 21 -nt dsRNA complementary to the promoter region of E-cadherin, p21, and VEGF (vascular endothelial growth factor) genes induced gene expression in a sequence-specific manner and dependency on Ago2, similar to RNAi. Janowski et al. (2007) Nat Chem Biol 3, 166-173 demonstrated an induced expression of progesterone receptor via complementary duplex RNAs targeting.

[0355] Aptamers are single-stranded synthetic DNA or RNA molecules, generally about 50 to about 150 nucleotides long, that can bind the nucleotide coding for proteins with high affinity and thus serve as decoys. DNA aptamers are short single-stranded oligonucleotide sequences similar to ASO with very high affinity for the target nucleic acids through structural recognition. DNA aptamers that target coding nucleotides for lysozyme, thrombin, human immunodeficiency virus trans-acting responsive element, hemin, interferon y, vascular endothelial growth factor, prostate specific antigen, dopamine and heat shock factor are under development. Aptamers can be isolated from a large pool of nucleic acids by a process called Systematic Evolution of Ligands by Exponential Enrichment (SELEX) or AptaBid.

[0356] Ribozymes can also function as inhibitors of expression for use in the present disclosure. Ribozymes are enzymatic RNA molecules capable of catalyzing the specific cleavage of RNA. The mechanism of ribozyme action involves sequence specific hybridization of the ribozyme molecule to complementary target RNA, followed by endonucleolytic cleavage. Engineered hairpin or hammerhead motif ribozyme molecules that specifically and efficiently catalyze endonucleolytic cleavage of variant transcript mRNA sequences are thereby useful within the scope of the present disclosure. Specific ribozyme cleavage sites within any potential RNA target are initially identified by scanning the target molecule for ribozymecleavage sites, which typically include the following sequences, GUA, GUU, and GUC. Once identified, short RNA sequences of between about 15 and 20 ribonucleotides corresponding to the region of the target gene containing the cleavage site can be evaluated for predicted structural features, such as secondary structure, that can render the oligonucleotide sequence unsuitable.

[0357] Strategies to inhibit the expression of oncogenes and the carcinogenic splice variants of essential genes at the mRNA level have been developed in the past few decades. Splicemodulating or splice-switching antisense oligonucleotides (SSOs) have been used for suppressing the genes involved in the progression of cancer (Dean N.M., Bennett C.F. Oncogene. 2003;22:9087-9096; Mercatante D.R., Mohler J.L., Kole R. J. Biol. Chem. 2002;277:49374-49382) or for blocking specific splice junctions to create novel isoforms that may serve as neoantigens (W02020 / 157760).

[0358] During the process of transcription, RNA polymerase converts genes into primary transcript mRNA (also known as pre-mRNA). This pre-mRNA usually contains introns, regions that will not go on to code for the final amino acid sequence. These introns are removed in the process of RNA splicing, leaving only exons, regions that will encode the protein. The resulting exon sequence constitutes mature mRNA, which is then read by the ribosome. Utilizing amino acids carried by transfer RNA (tRNA), the ribosome creates the polypeptide sequence in a process called translation.

[0359] RNA splicing is an essential process wherein precursor messenger RNA (pre-mRNA) is reshaped into mature mRNA. In alternative splicing, exons of any pre-mRNA get rearranged to form mRNA variants and subsequently protein isoforms, which are distinct both by structure and function. The process of splicing is catalyzed by the RNA-protein complex known as the spliceosome. During splicing, introns are removed, and exons are joined together.

[0360] SSOs act by binding to a pre-mRNA and disrupting the splicing of the gene transcript, e.g. any of the variant transcripts disclosed herein, by blocking the RNA-RNA base pairing or the protein-RNA binding interactions that occur between components of the splicing machinery (e.g. spliceosome).

[0361] SSOs can also act to increase splicing and expression of a transcript, thereby increasing the production of protein isoform.

[0362] A majority of the protein isoform identified that comprise non-exonic amino acid sequence are shortened variants due to premature stop codons. These variants are often missingimportant domains of the canonical protein. SSOs complementary to regions surrounding putative splicing sites of pre-mRNA can be used to “skip” the aberrant stop codon with relatively high efficiency, restore translation of the missing exons and therefore restore the missing activity.

[0363] SSOs can also be used to correct aberrant amino acid sequence due to mutations that result in a frameshifted ORF. SSOs complementary to regions surrounding putative splicing sites of pre-mRNA can be used to “skip” the aberrant sequence with relatively high efficiency, and therefore restore downstream canonical ORF. For example, the insertion of non-exonic sequence, e.g. TE sequence, may induce TE exonisation through creating a novel upstream start codon that produces additional exons or a frameshifted ORF. SSOs can restore the translation of the canonical isoform or fragments thereof.

[0364] SSOs can also be used to induce splicing events between coding exons and non-exonic sequence, e.g., TE sequence, and induce expression of tumor-specific antigen derived from a variant transcript disclosed herein, which comprises a junction between non-exonic sequence, e.g. TE sequence, and exonic sequence. Such SSOs can be complementary to a splice silencer site within a pre-mRNA sequence encoded by the variant transcript. In a particular embodiment, said complementary nucleic acid sequence is of 18 to 30 nucleotides in length. In another particular embodiment, said SSO comprises a chemically modified nucleic acid sequence, preferably a 2’-O-methyl (2’0Me) modified phosphorothioate oligonuleotide. In a particular embodiment, the non-exonic sequence is located at the 5 ’-end of the variant transcript and the exonic sequence is located at the 3 ’-end of the variant transcript or the non- exonic sequence is located at the 3’-end of the variant transcript and the exonic sequence is located at the 5 ’-end of the variant transcript.

[0365] In any of these SSO embodiments, the SSO may be complementary to a sequence within a region of about 40 to about 200 nucleotides upstream or downstream of the junction between non-exonic and exonic sequence, or within about 100 to about 500 nucleotides upstream or downstream, or within about 50, about 100, about 150, about 200, about 250, about 300, about 350, about 400, about 450, about 500, about 600, about 700, about 800, or about 900 nucleotides upstream or downstream of the junction. The SSO may be localized within the non- exonic sequence, e.g. the TE sequence, or the adjacent intronic sequence, or within the exonic sequence.Production and Delivery

[0366] A number of methods are conveniently used to deliver the oligonucleotide therapeutic agents described herein to the patient. For instance, the nucleic acid can be delivered directly, as “naked” DNA or RNA. The nucleic acids can also be delivered complexed to cationic compounds, such as cationic lipids. Delivery systems may optionally include cell-penetrating peptides, nanoparticulate encapsulation, virus like particles, liposomes, exosomes or any combination thereof. Dendrimers are a supermolecular delivery system which can be synthesized with various functional groups, making them a versatile non-viral particle delivery system. Polymers are also employed as delivery vehicles.

[0367] Oligonucleotide therapeutic agents for direct delivery can be prepared by known methods. These include techniques for chemical synthesis such as, e.g., by solid phase phosphoramadite chemical synthesis.

[0368] Alternatively, oligonucleotide therapeutic agents that are RNA can be generated by in vitro or in vivo transcription of DNA sequences encoding the RNA molecule. Such DNA sequences can be incorporated into a wide variety of vectors that incorporate suitable RNA polymerase promoters such as the T7 or SP6 polymerase promoters. Various modifications to any of the types of oligonucleotide therapeutic agents disclosed herein can be introduced as a means of increasing intracellular stability and half-life.

[0369] Oligonucleotide therapeutic agents may be delivered via a vector, as naked plasmid DNA, or as part of a viral vector. Vector constructs include plasmid, phage, transposon, cosmid, bacmid, mini-plasmids. Viral delivery vectors (e.g., viral particles) can be used to deliver the vector encoding the oligonucleotide therapeutic agent. Suitable viral vectors are known in the art and include adenovirus, parvoviruses including adeno-associated virus (AAV), poxvirus, papillomavirus, lentivirus, retrovirus, herpes virus, foamivirus, or Semliki Forest virus vector, including pseudotyped viruses. For example, AAV1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or 13 is a suitable virion. As another example, lentivirus pseudotyped with VSV is a suitable virion.Chemical modifications

[0370] In a particular embodiment, said oligonucleotide therapeutic agent comprises an oligonucleotide chemically modified. Chemical modification can be introduced at the backbone, nucleobase and / or sugar moiety. Possible modifications include but are not limited to the addition of flanking sequences of ribonucleotides or deoxyribonucleotides to the 5' and / or3' ends of the molecule, or alternating patterns of modifications, e.g. alternating 2’-0-Me and 2’-O-F modifications.

[0371] In one embodiment, the oligonucleotide therapeutic agent is modified by the substitution of at least one nucleotide with a modified nucleotide, such that in vivo stability is enhanced as compared to a corresponding unmodified oligonucleotide. In a related embodiment, the modified nucleotide is a sugar-modified nucleotide. In another embodiment, the modified nucleotide is a nucleobase-modified nucleotide.

[0372] In a particular embodiment, chemical modification is made to the phosphodiester backbone of said oligonucleotide to provide stability against nuclease degradation. For example, the non-bridging oxygen atom of the phosphate group is replaced with carbon (methyl phosphonate, phosphotriester), sulphur (phosphorothioate), nitrogen (phosphoroamidate) or boron (boranophosphate), preferably the oligonucleotide therapeutic agent is phosphorothioate oligonucleotide.

[0373] In another particular embodiment, the modified oligonucleotide therapeutic agent comprises a sugar-modified oligonucleotide to increase the binding affinity of oligonucleotides, protects oligonucleotide from nuclease degradation and increase its specificity. In particular, the modified nucleotide is a 2'-deoxy ribonucleotide. In certain embodiments, the 2'-deoxy ribonucleotide is 2'-deoxy adenosine or 2'-deoxy guanosine. In another embodiment, the modified nucleotide is a 2'-O-methyl (e.g., 2'-O-methylcytidine, 2'-O-methylpseudouridine, 2'- O-methylguanosine, 2'-O-methyluridine, 2'-O-methyladenosine, 2'-O-methyl)ribonucleotide, 2’ -O-m ethoxy ethyl ribonucleotide, locked nucleic acid (LNA), 2'-fluoro, 2'-amino, 2'-thio modified ribonucleotide, hexitol nucleic acid (HNA), cyclohexenyl nucleic acid (CeNA), altriol nucleic acid (ANA), 2’-O, 4’-C-ethylene bridged nucleic acid (ENA) or morpholino nucleic acid (MNA), preferably the SSO is a 2’-O-methyl oligonucleotide.

[0374] In a preferred embodiment, said oligonucleotide therapeutic agent comprises a 2’-O- methyl RNA phosphorothioate.Length of complementary sequence

[0375] Usually, according to the present disclosure, an oligonucleotide therapeutic agent comprises a complementary sequence from about 12 to about 30 nucleobases in length, or about 15 to about 20 nucleobases in length, or about 15 to about 25 nucleobases in length, or about 15 to about 30 nucleobases in length. Those skilled in the art appreciate that when affinityincreasing chemical modifications are used, the oligonucleotide therapeutic agent can be shorterand still retain specificity. Those skilled in the art will further appreciate that an upper limit on the size of the oligonucleotide therapeutic agent is imposed by the need to maintain specific recognition of the target sequence, and to avoid secondary structure forming self-hybridization of the oligonucleotide therapeutic agent and by the limitations of gaining cell entry. These limitations imply that an oligonucleotide therapeutic agent of increasing length (above and beyond a certain length which will depend on the affinity of the oligonucleotide therapeutic agent) will be more frequently found to be less specific, inactive or poorly active.

[0376] The oligonucleotide therapeutic agents according to the present disclosure may be made through the well-known technique of solid phase synthesis. Any other means for such synthesis known in the art may additionally or alternatively be used. It is well known to use similar techniques to prepare oligonucleotides such as the phosphorothioates and alkylated derivatives.Complementarity

[0377] The term “complementarity” with respect to oligonucleotide therapeutic agents refers to the capacity of base pairing, or hybridization, between the nucleobases of a first nucleic acid strand and the nucleobases of a second nucleic acid strand, mediated by hydrogen binding (e.g., Watson-Crick, Hoogsteen or reversed Hoogsteen hydrogen bonding) between corresponding nucleobases. The ability of the first and second nucleic acid strands to hybridize may be evaluated by the number of matches or mismatches between paired bases. For example, in DNA, adenine (A) matches / is complementary to thymine (T); and guanosine (G) is complementary to cytosine (C). For example, in RNA, adenine (A) is complementary to uracil (U); and guanosine (G) is complementary to cytosine (C). Some bases, e.g. inosine, are considered universal bases that pair with any other base. In certain embodiments, complementary nucleobase means a nucleobase of an antisense oligonucleotide that is capable of base pairing with a nucleobase of its target nucleic acid. For example, if a nucleobase at a certain position of an antisense oligonucleotide is capable of hydrogen bonding with a nucleobase at a certain position of a target nucleic acid, then the position of hydrogen bonding between the oligonucleotide and the target nucleic acid is considered to be complementary at that nucleobase pair. Nucleobases comprising certain modifications may maintain the ability to pair with a counterpart nucleobase and thus, are still capable of nucleobase complementarity. Alternatively, the ability of the first and second nucleic acid strands to hybridize may be evaluated under stringent conditions such as 400 mM NaCl, 40 mM PIPES pH 6.4, 1 mM EDTA, 50°C or 70°C for 12- 16 hours followed by washing (see, e.g., "Molecular Cloning: ALaboratory Manual, Sambrook, et al. (1989) Cold Spring Harbor Laboratory Press). Non- Watson-Crick base pairing can also occur through other hydrogen bond interactions, or interactions between C-H and O / N groups, such as Hoogsteen A:U and G:U wobble pairs. The first and second nucleic acid strands may hybridize sufficiently to each other to modulate expression when they are less than 100% complementarity. Thus complementary as used herein includes base pairing that is 100% complementary, or about 95%, about 90%, about 85%, about 80%, about 75%, or about 70% complementary, as long as it is sufficient to modulate gene expression.

[0378] As used herein, modulation of gene expression includes any increase or decrease in expression or protein activity or level of the gene of interest, or encoded protein, as compared to a situation wherein no modulation has been induced. The difference can be of at least, 20 %, 30 %, 40 %, 50 %, 60 %, 70 %, 80 %, 90 %, 95 %, 99 % as compared to the normal expression of the gene or level of the protein which has not been targeted.Pharmaceutical compositions

[0379] Pharmaceutically acceptable carriers typically enhance or stabilize the composition, and / or can be used to facilitate preparation of the composition. Pharmaceutically acceptable carriers include solvents, dispersion media, coatings, antibacterial and antifungal agents, isotonic and absorption delaying agents, and the like that are physiologically compatible and in some embodiments pharmaceutically inert.

[0380] Administration of a pharmaceutical composition comprising antibodies as herein disclosed, including antigen-binding fragments, can be accomplished orally or parenterally. Methods of parenteral delivery include topical, intra-arterial (directly to the tumor), intramuscular, spinal, subcutaneous, intramedullary, intrathecal, intraventricular, intravenous, intraperitoneal, or intranasal administration.

[0381] Thus, in addition to the active ingredients, these pharmaceutical compositions may contain suitable pharmaceutically acceptable carriers comprising excipients and auxiliaries which facilitate processing of the active compounds into preparations which can be used pharmaceutically. Further details on techniques for formulation and administration may be found in the latest edition of Remington's Pharmaceutical Sciences (Ed. Maack Publishing Co, Easton, Pa.).

[0382] Depending on the route of administration, the active compound, i.e., antibody, bispecific and multispecific molecule, may be coated in a material to protect the compound from the action of acids and other natural conditions that may inactivate the compound.

[0383] The composition is typically sterile and preferably fluid. Proper fluidity can be maintained, for example, by use of coating such as lecithin, by maintenance of required particle size in the case of dispersion and by use of surfactants. In many cases, it is preferable to include isotonic agents, for example, sugars, polyalcohols such as mannitol or sorbitol, and sodium chloride in the composition. Long-term absorption of the injectable compositions can be brought about by including in the composition an agent which delays absorption, for example, aluminum monostearate or gelatin.

[0384] Pharmaceutical compositions for oral administration can be formulated using pharmaceutically acceptable carriers well known in the art in dosages suitable for oral administration. Such carriers enable the pharmaceutical compositions to be formulated as tablets, pills, dragees, capsules, liquids, gels, syrups, slurries, suspensions and the like, for ingestion by the patient.

[0385] Pharmaceutical preparations for oral use can be obtained through combination of active compounds with solid excipient, optionally grinding a resulting mixture, and processing the mixture of granules, after adding suitable auxiliaries, if desired, to obtain tablets or dragee cores. Suitable excipients are carbohydrate or protein fillers such as sugars, including lactose, sucrose, mannitol, or sorbitol; starch from corn, wheat, rice, potato, or other plants; cellulose such as methyl, cellulose, hydroxypropylmethylcellulose, or sodium carboxymethylcellulose; and gums including arabic and tragacanth; and proteins such as gelatin and collagen. If desired, disintegrating or solubilizing agents may be added, such as the cross-linked polyvinyl pyrrolidone, agar, alginic acid, or a salt thereof, such as sodium alginate.

[0386] Dragee cores are provided with suitable coatings such as concentrated sugar solutions, which may also contain gum arabic, talc, polyvinylpyrrolidone, carbopol gel, polyethylene glycol and / or titanium dioxide, lacquer solutions, and suitable organic solvents or solvent mixtures. Dyestuffs or pigments may be added to the tablets or dragee coatings for product identification or to characterize the quantity of active compound, ie. dosage.

[0387] Pharmaceutical preparations that can be used orally include push-fit capsules made of gelatin, as well as soft, sealed capsules made of gelatin and a coating such as glycerol or sorbitol. Push-fit capsules can contain active ingredients mixed with a filler or binders such aslactose or starches, lubricants such as talc or magnesium stearate, and optionally, stabilizers. In soft capsules, the active compounds may be dissolved or suspended in suitable liquids, such as fatty oils, liquid paraffin, or liquid polyethylene glycol with or without stabilizers.

[0388] Pharmaceutical formulations for parenteral administration include aqueous solutions of active compounds. For injection, the pharmaceutical compositions of the invention may be formulated in aqueous solutions, preferably in physiologically compatible buffers such as Hank's solution, Ringer's solution, or physiologically buffered saline. Aqueous injection suspensions may contain substances that increase viscosity of the suspension, such as sodium carboxymethyl cellulose, sorbitol, or dextran. Additionally, suspensions of the active compounds may be prepared as appropriate oily injection suspensions. Suitable lipophilic solvents or vehicles include fatty oils such as sesame oil, or synthetic fatty acid esters, such as ethyl oleate or triglycerides, or liposomes. Optionally, the suspension may also contain suitable stabilizers or agents which increase the solubility of the compounds to allow for the preparation of highly concentrated solutions.

[0389] For topical or nasal administration, penetrants appropriate to the particular barrier to be permeated are used in the formulation. Such penetrants are generally known in the art.

[0390] Pharmaceutical compositions of the disclosure can be prepared in accordance with methods well known and routinely practiced in the art. See. e.g., Remington: The Science and Practice of Pharmacy, Mack Publishing Co., 20th ed., 2000; and Sustained and Controlled Release Drug Delivery Systems, J R. Robinson, ed., Marcel Dekker, Inc., New York, 1978. Pharmaceutical compositions are preferably manufactured under GMP conditions.Therapeutic and diagnostic methods

[0391] In any of the embodiments, the products described herein may be used in methods for treating a disease associated with a protein isoform corresponding to a gene function described in any of Tables 1 A, 3, or 4. Such uses include uses for the products including novel protein isoforms, fragments thereof, and polynucleotides encoding such protein isoforms or fragments thereof; neoantigenic peptides of the novel protein isoforms; polynucleotides encoding such protein isoforms, fragments, or peptides, optionally linked to one or more heterologous regulatory control nucleotide sequences; vectors comprising such polynucleotides; oligonucleotide therapeutic agents targeting such polynucleotides; vaccine or immunogenic compositions comprising such protein isoforms or fragments thereof or neoantigenic peptides thereof, or such polynucleotides; an antibody, or an antigen-binding fragment thereof, a T cellreceptor (TCR) in particular a non-HLA restricted TCR, or a chimeric antigen receptor (CAR) that specifically binds such protein isoforms or fragments thereof; methods of producing such antibodies, TCRs or CARs; polynucleotides encoding such antibodies, CARs or TCRs, optionally linked to one or more heterologous regulatory control nucleotide sequences; host cells, dendritic cells, antigen-presenting cells (APC) or immune cells that specifically bind to such protein isoforms or fragments thereof described herein. In some embodiments, particularly for the protein isoforms in Tables 3 or 4, the use is for inhibiting proliferation of cancer cells, or for the treatment of cancer, in patients suffering from cancer, or for the prophylactic treatment of cancer, in patients at risk of cancer.

[0392] Cancers that can be treated using the therapy described herein include any solid or non-solid tumors as previously defined. Of particular interest according to the present disclosure are breast cancer, melanoma and lung cancer.

[0393] Cancers includes also the cancers which are refractory to treatment with other chemotherapeutics. The term “refractory, as used herein refers to a cancer (and / or metastases thereof), which shows no or only weak antiproliferative response (e.g., no or only weak inhibition of tumor growth) after treatment with another chemotherapeutic agent. These are cancers that cannot be treated satisfactorily with other chemotherapeutics. Refractory cancers encompass not only (i) cancers where one or more chemotherapeutics have already failed during treatment of a patient, but also (ii) cancers that can be shown to be refractory by other means, e.g., biopsy and culture in the presence of chemotherapeutics.

[0394] The therapy described herein is also applicable to the treatment of patients in need thereof who have not been previously treated.

[0395] A subject as per the present disclosure is typically a patient in need thereof. In some embodiments, the subject has been diagnosed with cancer or is at risk of developing cancer. The subject is typically a human, dog, cat, horse or any animal in which a protein isoform- specific immune response is desired.

[0396] The present disclosure also pertains to uses of any of the novel protein isoforms, or fragments thereof or neoantigenic peptides thereof; polynucleotides or vectors encoding such proteins, fragments or peptides thereof; oligonucleotide therapeutic agents; vaccine or immunogenic compositions; host cells or dendritic cells or APCs; antibody, antigen-binding fragment thereof, TCR or CAR; or an immune cell as described herein; for use in vaccination therapy, e.g. cancer vaccination therapy of a subject, or for treating a disease associated withthe gene function associated with the protein isoform in any of Tables 1 A, 3 or 4, or for treating cancer in a subject. In some embodiments, the peptide(s) binds at least one MHC molecule of said subject.

[0397] The present disclosure also provides a method for treating a disease associated with the gene function associated with the protein isoform in any of Tables 1A, 3 or 4, or for treating cancer in a subject comprising administering a vaccine or immunogenic composition as described herein to said subject in a therapeutically effective amount to treat the subject. The method may additionally comprise the step of identifying a subject who has the disease, e.g., cancer.

[0398] The present disclosure also relates to a method of treating a disease associated with the gene function associated with the protein isoform in any of Tables 1A, 3 or 4, or for treating cancer in a subject comprising producing an antibody or antigen-binding fragment thereof, or a polynucleotide encoding such antibody or antigen-binding fragment thereof, or an immune cell comprising said antibody or antigen-binding fragment thereof or encoding polynucleotide, by the method as herein described and administering to a subject having said disease said antibody or antigen-binding fragment thereof, or encoding polynucleotide, or said immune cell, in a therapeutically effective amount to treat said subject.

[0399] The present disclosure also relates to an antibody (including variants and derivatives thereof), a T cell receptor (TCR) (including variants and derivatives thereof), a non-HLA restricted TCR (HI TCR), or a CAR (including variants and derivatives thereof) which are directed against a protein isoform as herein described, or neoantigenic peptide typically in association with an MHC or HLA molecule, for use in therapy of a subject, for treating a disease associated with the gene function associated with the protein isoform in any of Tables 1A, 3 or 4, or for treating cancer in a subject.

[0400] In some embodiments said antibody, TCR (in particular non-HLA restricted TCR) or CAR specifically binds a protein isoform as herein defined, e.g. comprising any of the amino acid sequences of Table 1 A. Typically said antibody, TCR (in particular non-HLA restricted TCR) or CAR comprise an antigen binding fragment or domain as previously defined.

[0401] The present disclosure also relates to an antibody (including antigen-binding fragments, variants and derivatives thereof), a T cell receptor (TCR) (including variants and derivatives thereof), or a CAR (including variants and derivatives thereof) which are directed against a protein isoform, fragment thereof, or neoantigenic peptide thereof (typically inassociation with an MHC or an HLA molecule) as herein described, or a cell, including an immune cell, which targets a protein isoform, fragment thereof, or neoantigenic peptide thereof, e.g. a cell, including an immune cell, comprising said antibody, TCR or CAR, or encoding polynucleotide, as previously defined, for use in adoptive cell or CAR-T cell therapy in a subject. In some embodiments said antibody, TCR (in particular non-HLA restricted TCR) or CAR binds a protein isoform as herein defined and notably a protein isoform of any of the amino acid sequences of Table 1A. Typically said antibody, TCR (in particular non-HLA restricted TCR) or CAR binds comprise an antigen binding fragment or domain (binding a protein isoform of any of the amino acid sequences of Table 1 A) as previously defined. Thus typically in some embodiments the immune cell targets a protein isoform as herein defined. Typically, the skilled person is able to select an appropriate antigen receptor which binds and recognizes a neoantigenic peptide as previously defined with which to redirect an immune cell to be used for use in cancer cell therapy. In a particular embodiment, the immune cell for use in the method of the present disclosure is a redirected T-cell, e.g. a redirected CD8+ and / or CD4+ T-cell.

[0402] In some embodiments, cancer treatment, vaccination therapy and / or adoptive cell cancer therapy as above described are administered in combination with additional therapies suitable for treating the disease, e.g. cancer therapies. In particular, the T cell compositions according to the present disclosure may be administered in combination with checkpoint blockade therapy, co-stimulatory antibodies, chemotherapy and / or radiotherapy, targeted therapy or monoclonal antibody therapy.

[0403] Checkpoint inhibitors include, but are not limited to, PD-1 inhibitors, PD-L1 inhibitors, Lag-3 inhibitors, Tim-3 inhibitors, TIGIT inhibitors, BTLA inhibitors, V-domain Ig suppressor of T-cell activation (VISTA) inhibitors and CTLA-4 inhibitors, IDO inhibitors for example. Co-stimulatory antibodies deliver positive signals through immune-regulatory receptors including but not limited to ICOS, CD137, CD27 OX-40 and GITR. In a preferred embodiment the checkpoint inhibitor is a CTLA-4 inhibitor.

[0404] A chemotherapeutic entity as used herein refers to an entity which is destructive to a cell, that is the entity reduces the viability of the cell. The chemotherapeutic entity may be a cytotoxic drug. A chemotherapeutic agent contemplated includes, without limitation, alkylating agents, anthracyclines, epothilones, nitrosoureas, ethylenimines / methylmelamine, alkyl sulfonates, alkylating agents, antimetabolites, pyrimidine analogs, epipodophylotoxins, enzymes such as L-asparaginase; biological response modifiers such as IFNa, IL-2, G-CSF andGM-CSF; platinum coordination complexes such as cisplatin, oxaliplatin and carboplatin, anthracenediones, substituted urea such as hydroxyurea, methylhydrazine derivatives including N-methylhydrazine (MIH) and procarbazine, adrenocortical suppressants such as mitotane (o,p'-DDD) and aminoglutethimide; hormones and antagonists including adrenocorticosteroid antagonists such as prednisone and equivalents, dexamethasone and aminoglutethimide; progestin such as hydroxyprogesterone caproate, medroxyprogesterone acetate and megestrol acetate; estrogen such as diethylstilbestrol and ethinyl estradiol equivalents; antiestrogen such as tamoxifen; androgens including testosterone propionate and fluoxymesterone / equivalents; antiandrogens such as flutamide, gonadotropin-releasing hormone analogs and leuprolide; and non-steroidal antiandrogens such as flutamide.

[0405] 'In combination' may refer to administration of the additional therapy before, at the same time as or after administration of the immune cell composition according to the present disclosure.

[0406] In addition or as an alternative to the combination with checkpoint blockade, the immune cell composition of the present disclosure may also be genetically modified to render them resistant to immune-checkpoints using gene-editing technologies including but not limited to TALEN and Crispr / Cas. Such methods are known in the art, see e.g. US20140120622. Gene editing technologies may be used to prevent the expression of immune checkpoints expressed by immune cells including but not limited to PD-1 , Lag-3, Tim-3, TIGIT, BTLA CTLA-4 and combinations of these. The immune cell as discussed here may be modified by any of these methods.

[0407] The immune cell according to the present disclosure may also be genetically modified to express molecules increasing homing into tumours and or to deliver inflammatory mediators into the tumour microenvironment, including but not limited to cytokines, soluble immune- regulatory receptors and / or ligands.

[0408] The present disclosure also provides diagnostic methods, for example, a method for determining the prognosis of a cancer or for detecting cancer cells characterized by increased expression of a protein isoform of Table 3 or 4.

[0409] In some embodiments, the method comprises the step of detecting or quantifying a protein isoform described herein, e.g. a protein isoform of any of the amino acid sequences of Table 1A, in a biological sample from a patient. Such methods may involve contacting the biological sample with an antibody, including antigen-binding fragment thereof, thatpreferentially binds to the protein isoform compared to the canonical protein, e.g., an antibody specific for the non-exonic amino acid sequence of said protein isoform. The methods may also involve detecting binding of the antibody with the protein isoform.

[0410] In some embodiments, the method comprises the step of detecting or quantifying a nucleic acid encoding any of the amino acid sequences of Table 1A in a biological sample isolated from a patient. In some embodiments, the detecting comprises: (a) contacting said nucleic acid from the biological sample with a probe which specifically hybridizes to any of the nucleotide sequences of Table 1A, preferably a probe that hybridizes to the junction of exonic sequence and non-exonic sequence; (b) optionally detecting a complex formed between the probe and the nucleic acid of any of the nucleotide sequences of Table 1A from the biological sample, and comparing the amount of the complex in the biological sample to the amount of the complex in a reference sample comprising corresponding normal tissue; (c) wherein detecting said complex in the biological sample is associated with a better or worse prognosis of cancer, (d) or wherein detecting an increased amount of said complex in the biological sample as compared to the corresponding normal tissue indicates that the biological sample comprises cancer cells. In particular, the protein isoforms of Table 4 are associated with cancer, and the protein isoforms of Table 4 are associated positively or negatively with cancer survival.

[0411] Said antibodies or probes for detecting or quantifying are preferably labeled with, e.g. a chemiluminescent, radioactive, magnetic, nanoparticle, or other label. Said biological sample includes a sample of tissue, plasma, blood, or other fluid from the patient.EXAMPLES1. Example 1: Identification of novel and stably expressed splice variants containing TE / exon junction

[0412] A combination of transcriptome assembly, ribosome profiling, and mass spectrometry was applied to identify 1300 unannotated protein isoforms generated by non-canonical splicing between exons and transposable elements (TE). See Figure 1. First, RNA transcripts containing splice junctions between exons and TEs (JETs) were identified by RNAseq mapping to both a protein-coding exon and TE in tumor samples from The Cancer Genome Atlas (TCGA, n=9191) and the Cancer Cell Line Encyclopedia (CCLE), generally as described in Merlotti et al. (2023), Sci. Immunol. 8, eabm6359 and Burbage et al. (2023), Sci. Immunol. 8, eabm6360.1.1 Identification of recurrent JETs in TCGA and CCLE datasets

[0413] To identify splicing junctions between exons and TEs, we used a bioinformatic pipeline that detects spliced RNAseq reads mapping to both a protein-coding exon and a TE, as in Merlotti A. et al. (23). The pipeline was applied to all tumor samples from The Cancer Genome Atlas (TCGA, n=9191). Two examples of identified JETs are illustrated in Figure 2. In total, 22,947 JETs are detected in at least 2 patients. The number of JETs per sample ranges from 6 to 512, with an average of 236 JETs per sample (Figure 3). Even though most JETs are not recurrent, 2506 JETs are shared by at least 1% of the samples, and 467 by at least 10% (Figure 4). JET patient recurrence distribution is similar in all TCGA cancer types (examples in LU AD and LGG are shown in Figure 5). Unsupervised clustering of JETs expressed in over 10% of TCGA samples reveals JETs preferentially expressed in certain cancer types, and JETs expressed across all indications. UMAP visualization based on JET expression shows clustering of TCGA samples according to the tissue of origin (Figure 6). Recurrent JETs are also present in tumor-adjacent normal tissues from TCGA (n=679, Figure 7). JET recurrence in TCGA tumors is highly correlated with recurrence in normal tissues (R2=0.88, Figure 8), indicating that even while some JETs are tumor-specific, most are found in both tumor and healthy tissues. Furthermore, to validate recurrent JETs in an independent cohort, we used the Cancer Cell Line Encyclopedia (CCLE), which contains RNAseq data from 1019 cell lines. The overlap between datasets increases proportionally with JET recurrence. In total, 12,953 JETs are recurrent in at least 1% of the samples from TCGA and / or CCLE. We conclude that a subpopulation of JETs is recurrent across individuals and independent cohorts and presents tissue-dependent expression profiles.1.2 Relative expression levels of JETs compared to canonical exon-exon junctions

[0414] All recurrent JETs, defined as transcripts with (a) level of expression over 2*10'7and (b) present in more than 1% of tumor TCGA and / or CCLE, were analyzed. 2506 JETs are shared by at least 1% of the samples in TCGA, and 467 by at least 10% of samples. Unsupervised clustering of JETs expressed in over 10% of TCGA samples reveals some JETs preferentially expressed in certain cancer types, and other JETs expressed across all indications.

[0415] To investigate the levels of expression of these novel splicing junctions, the expression of recurrent JETs was compared with expression of the corresponding canonical exon-exon junctions (according to Gencode annotation). To calculate the proportion of a junction among all overlapping splicing events, the CPM expression value was divided against the CPM values of all junctions involving the same breakpoint of the canonical exon. All junctions were considered; and no thresholds of expression were used. On average, recurrent JETs have 10-fold lower expression levels than canonical junctions. JET and canonical junction expression levels for the same genes are not correlated. 572 JETs, however, contribute to over 50% of all splicing junctions for the corresponding exon in at least one TCGA indication. Interestingly, JETs that represent a few percent of all splicing events for a given exon can be as recurrent in patients as JETs representing much higher proportions among all splicing events. The analysis of non-canonical splicing junctions between exons and TEs revealed a population of low abundance, unannotated splicing isoforms of exonized TEs that can be highly recurrent across individuals.1.3 Identification of JETs that are translated and encode unannotated protein isoforms, through genome-guided transcriptome assembly and ribosome profiling analysis

[0416] RNAseq data was processed using StringTie and correlated with RiboSeq data as follows. Raw RNA-seq files were aligned using STAR (v2.5.3a) single-pass and two-pass modes. The outputs of both modes were processed in downstream processes in parallel. Bam files for each cell line were processed by StringTie v2.1.4. A consensus gtf file was then generated using the StringTie-merge option, and it was merged with the hgl9 human reference genome (Ensembl).

[0417] RiboSeq was performed by a commercial service on H1650 and H1395 cell lines. Additional RiboSeq publicly available data was obtained from Calviello et al. (2020), Nat. Struct. Mol. Biol. 27, 717-725; Martinez et al. (2020), Nat. Chem. Biol. 16, 458-468; Clamer et al., Cell Rep. 25, 1097-1108. e5; Park et al. (2016), Mol. Cell 62, 462-471; Calviello et al. (2016), Nat. Methods 13, 165-170.

[0418] Adaptors from RiboSeq FASTQ were removed using cutadapt vl.8 68, and rRNA contaminants were discarded by aligning using bowtie2 v2.2.5. Unaligned reads were then mapped against the assembled transcriptome using STAR (Spliced Transcripts Alignment to a Reference). RiboseQC vl. l was used to deduce P-sites. Translation of transcripts containing at least one JET recurrent more than 1% of TCGA / CCLE was then interrogated using ORF quant vl.02.0. The called ORFs were blasted against RefSeq Curated protein databases (retrieved on December 2022). ORFs that overlapped annotated CDS with less than 95% similarity and / or with the insertion of 5 amino acids were considered as unannotated isoforms. Among the 12,953 JETs expressed in at least in 1% of TCGA and / or CCLE samples, 3,801 are assembled into transcripts using in-house and publicly available RNAseq datasets and 1,292 JETs are unannotated ORFs (JET-ORFs), among which partially overlapping annotated proteinsequences were identified, corresponding to 820 unique translated JETs recurrent in TCGA. While the TEs located in the start (5’ end) and internal exons of the ORF represent 12% and 14% of the JET-ORFs, respectively; over 70% of the identified ORFs (950 JET-ORFs) contain the TE in the 3’ end (and introduce a new stop codon). See Figure 9 and Figure 10. Consistent with this observation, JET-ORFs are overall shorter than the corresponding canonical ORFs. Table 1 below shows the JET-ORF amino acid sequence and associated variant transcript nucleotide sequence, including nucleotide sequence upstream and downstream of the proteincoding sequence which may comprise regulatory control nucleotide sequences for the variant. Table 1 also shows the gene name of the canonical gene from which the protein isoform is derived; the genomic coordinates of the variant transcript; its chimeric ID; gene expression level; expression level of the variant transcript in tumor tissue; expression level of the variant transcript in normal tissue; and relative expression of the variant transcript in the GTex database shows that the transcript is tumor-associated in a certain percentage of subjects with cancer (TA with filters, without filters, and excluding testis).TABLE 1AIll

[0419] JET-ORFs of Table 1A were analyzed for the presence of transmembrane domains, indicating that they are expressed on the cell surface. The predicted transmembrane domains and their topology are shown in Table IB below.TABLE IB

[0420] The expression levels of JET-ORFs were compared to expression levels of canonical ORFs, to investigate if JET-ORFs are translated as efficiently. RNAseq and RiboSeq levels in H1395 and H1650 cell lines were analyzed (for which coupled RNAseq / RiboSeq datasets were generated). Some JETs were detected at higher levels in RiboSeq compared to RNAseq, suggesting differences in the efficiency of translation. Translation efficiency was defined by the ratio between RiboSeq CPM expression and RNAseq CPM expression. Translated JETs have on average lower expression levels than canonical ORFs in both RNAseq and RiboSeq, but display slightly higher overall translation efficiencies compared to canonical ORFs. To confirm that JETs are efficiently translated, their susceptibility to nonsense-mediated mRNA decay (NMD) was assessed. NMD is a surveillance mechanism that induces the degradation of mRNA transcripts with premature termination sites after a pioneer round of translation, thereby preventing the production of deleterious protein products. We treated 4 cell lines with puromycin to inhibit NMD, and performed RNAseq (Figure 1). Most JETs (84%) are not differentially expressed between conditions, indicating that they are not susceptible to NMD. These results show that JETs can be efficiently translated and represent a source of non- canonical protein isoforms.

[0421] The instability index of JET-ORF was calculated using the instalndex function from the Peptides R package. JET-ORFs have overall the same predicted stability as canonical ORFs.

[0422] The population of TE elements (LINEs, SINEs, LTR and DNA transposons) was also analyzed. In comparison to transcriptome (i.e., assembled JET-containing transcripts), translated JETs (i.e., detected by RiboSeq) are enriched in LINEs (adjusted p. value <0.0001). Overall, LINEs are the most enriched TEs in translated JET-ORFs that appear to be novel protein isoforms, and exonized LINEs appear to be preferentially translated over SINE- containing transcripts.1.4 Confirmation of JET expression using mass spectrometry -based proteomics

[0423] Publicly available mass spectrometry proteomics was conducted in 6 cell lines. Data was obtained from MASSIVE repository (MSV000086944) and processed using MSFragger v3.7 in FragPipe vl9.1 environment with the following parameters: precursor mass tolerancelOppm and fragment mass tolerance 0.02 Da. Methionine oxidation (+15.995Da), N-acetylation (+42.01 IDa) were enabled as dynamic modifications. Carbamidomethylation (+57.021Da) was considered as fixed modifications. Enzymatic digestion was selected accordingly (trypsin, chymotrypsin, AspN, GluC, LysC and LysN), achieving high coverage of the proteoform diversity. MSBooster rescoring was enabled and Percolator was used to filter at a false discovery rate (FDR) of 1% at peptide level. No FDR was used at the protein level. MS / MS spectra were searched against the human proteome from Uniprot / SwissProt with isoforms (updated 06.03.2020) and concatenated with the JET-ORFs. Identified peptides were filtered by human proteome (SwissProt + RefSeq) considering L and I as equivalent. Raw files from HeLaS3 and K562 cell lines were run in Comet in similar parameters as a validation method.

[0424] 170 peptides mapped uniquely to a JET-ORF, derived from 119 different JET-ORFs identified through MS-proteomics that correspond to the indicated chromosomal coordinates, and which are isoforms of the gene listed. These results provide additional evidence that JETs represent non-canonical protein isoforms.2. Example 2: Correlation with cancer patient survival

[0425] Correlation of JET-ORF expression with patient survival was analyzed in 31 tumor indications in TCGA). Survminer package was used to calculate and represent survival curves based on JET expression in TCGA. When the JET of interest was expressed in more than half of the samples, the comparison groups were organized according to “above or below” the median JET expression. When the JET was expressed in less than half of the samples, the absence or presence of the JET was used to divide comparison groups. Bonferroni p-value adjustment was performed based on the total number of JETs interrogated in each indication.

[0426] From the initial 1292 JET-ORFs from recurrent JETs in TCGA, 201 JETs are correlated with patient survival in at least 1 clinical cohort tumor type. After Bonferroni correction for adjusted p values, 23 JETs remain significantly positively or negatively correlated with survival in at least 1 of the two indications. As one example, a JET in HIBADH gene, a mitochondrial dehydrogenase, is negatively correlated with the survival probability in TCGA-LUAD, while no significant differences are observed with the canonical junction. The overexpression of a JET in the BRK1 gene, and not its canonical junction, is associated with an increased survival in LGG. See examples in Table 3 below. The adjusted P value correction shows significance as indicated in the Bonferroni column. The ALDH3A2 JET-ORF ispositively correlated with survival in KIRC, MESO, and OV tumor types, while the canonical junction is not correlated.

[0427] TABLE 3

[0428] Taken together, the data lead to the conclusions that i) non-canonical splicing isoforms can be recurrent and contribute to the cell proteome with a large population of lowly abundant, but potentially functional protein isoforms, ii) these isoforms appear when relatively recent TEs (preferentially LINEs) are exonized (mainly from introns in ancient genes, and iii) JET expression can be correlated with cancer patient survival, suggesting that may be biologically relevant. The non-canonical isoforms derived from TE exonization represent functional protein variants that may vary in cellular localization and function.3. Example 3: JET-derived novel isoforms from genes described as cancer drivers

[0429] Among the 1292 identified JET-ORFs, 64 were derived from a gene annotated as cancer driver, tumor suppressor, or oncogene by the following sources: IntoGene, COSMIC, and TSGene databases. The list of the JET-containing cancer driver genes and the database where it is annotated is shown in Table 4.TABLE 44. Example 4: Cellular localization and function of JETs expressed using mNeonGreen split fluorescence system

[0430] Five JET isoforms, H2AFY, ALDH3A2, IL15RA, PTEN and WWOX, were selected for ectopic expression in HeLa cells using mNeonGreen (mNG) split fluorescence system. JET- ORFs were tagged with the mNGl 1 fragment, and complementation with the remaining mNGl- 10 constitutively expressed by host cells (HeLa mNGl-10) was detected by flow cytometry or microscopy.

[0431] Expression plasmids were produced encoding mNGl 1 tagged ORFs under control of a SFFV promoter. Lentivirus particles were produced by HEK293T-Lenti-X cell lines (Takara) transfected with the plasmid containing the mNGl l tagged JET-ORF, together with envelope (pVSVG) and packaging (psPAX2) plasmids. After 60 h, supernatant was collected, and lentiviral particles were ultracentrifugated (31000 g) in a 20% sucrose gradient. Lentivirus was aliquoted and stocked at -80°C. For transduction, 50uL of lentivirus suspension was transferred to 0.25M HeLa or HEK293FT cells expressing mNG 1-10. Ectopic expression was evaluated at least after 48h by flow cytometry and confocal microscopy.

[0432] For confocal microscopy, target cells were plated in 10mm petri dishes at a concentration of 50,000 cells / mL in DMEM 10% FBS. After overnight incubation, the medium was replaced with culture medium containing 1 / 5000 dilution of sirDNA (Spirochrome), and it was incubated for 30-60 minutes. Then, medium was replaced with FluoroBrite DMEM medium (Gibco) and live imaging was performed using Inverted Eclipse Ti-E (Nikon) and Spinning disk CSU-Xlmiscroscope (Yokogawa) integrated in Metamorph software (Gataca Systems). Fluorescence intensity was quantified using Image J 71.

[0433] The H2AFY histone (also called macroH2A) JET-ORF, like the canonical isoform, localized mainly to the nucleus and colocalized with chromatin in both interphase and mitosis. The ALDH3A2 JET-ORF localized to the endoplasmic reticulum, like the canonical isoform, an aldehyde dehydrogenase. TE exonization did not impair the ALDH3A2 transmembrane domain, according to TMHMM predictor, and purification of membranes from ALDH3A2 JET-ORF expressing cells confirmed the transmembrane insertion. The IL15RA JET-ORF also preserves the transmembrane domain and is expressed at the plasma membrane. The PTEN and WWOX JET-ORFs (2 tumor suppressors), in contrast, present different subcellular locations compared to the corresponding canonical ORFs. While the PTEN canonical ORF is mainly cytosolic and the WWOX canonical ORF localizes to the Golgi, the 2 JET-ORFs present clearnuclear enrichment compared to the canonical ORF. Quantification of the mean fluorescence ratio between the nucleus and the total cell confirms the preferential nuclear localization of PTEN and WWOX JET-ORFs. Taken together, the data show that JET-ORFs can be stable and localize to specific subcellular locations.

[0434] The PTEN JET-ORF and WWOX JET-ORF not only displayed altered subcellular localization, but also displayed altered function. To analyze possible divergent functions between the JET and canonical isoforms, the differential expression of genes in cells after overexpression of either the JET-ORF or the canonical ORF was compared. Functional association network analysis was also performed on the genes commonly upregulated by both isoforms such as ATF3 and PDGF. PTEN JET-ORF induces the overexpression of a gene network related to IL-6, suggesting that it activates inflammatory pathways, unlike the canonical ORF. WWOX JET-ORF activates different transcription factors compared to the canonical ORF. For example, while the canonical ORF induces an ATF6 response (a transcription factor located in the endoplasmic reticulum), WWOX JET-ORF activates KLF5, TWIST 1 and TP73 transcription factors.

Claims

CLAIMS1. A method for identifying a novel protein isoform, wherein the novel protein isoform comprises non-exonic amino acid sequence, comprising the steps of:(a) aligning RNA sequences reads from a mammalian cell or tissue to a reference mammalian genome,(b) assembling transcripts from the genome-aligned RNA sequence reads of step (a), optionally using the combination of STAR and StringTie,(c) selecting variant transcripts from step (b) that (i) comprise one or more mutations compared to a reference mammalian genome, (ii) comprise exonic sequence, and (iii) comprise non-exonic sequence, optionally wherein the non-exonic sequence is a transposable element (TE),(d) optionally selecting variant transcripts from step (b) or step (c) that contain junctions between exonic sequence and non-exonic sequence and optionally are recurrently expressed in at least 1 % of subjects from a population of subjects suffering from a disease, optionally cancer, and(e) identifying the open reading frames (ORFs) of the variant transcripts from step (c) or (d) that are translated into protein, thereby identifying the novel protein isoform.

2. The method of claim 1 further comprising the step of selecting variant transcripts present in mRNA subpopulations of a mammalian cell or tissue identified as directly bound to ribosomes, optionally using RiboSeq sequencing, RiboseQC and / or ORFquant; and optionally comprising the step of selecting the ORFs from step (e) that are present in peptide fragments identified through mass spectrometry proteomics.

3. The method of any of claims 1-2 wherein (a) the non-exonic region of the reference mammalian genome is a TE, optionally intronic TE, optionally a short interspersed nuclear element (SINE), or a long interspersed nuclear element (LINE), or a long terminal repeat (LTRs), and / or (b) the exonic sequence is from a gene of phylostratum 1.

4. An isolated protein isoform that comprises an amino acid sequence encoded by a variant transcript produced by the method of any of claims 1-3, or a fragment thereof at least 4, 5, 6, 7 or 8 amino acids in length.

5. An isolated protein isoform that comprises non-exonic amino acid sequence, optionally (a) that overlaps a junction of exonic amino acid sequence and non-exonic amino acid sequence, or (b) comprises sequence encoded by a portion of a TE, or (c) is encoded by a frameshifted ORF downstream of the junction of exonic sequence and non-exonic sequence; or a fragment thereof at least 4, 5, 6, 7 or 8 amino acids in length.

6. An isolated protein isoform comprising any of the amino acid sequences of Table 1 A, or a fragment thereof at least 4, 5, 6, 7 or 8 amino acids in length.

7. A peptide that is a fragment about 8 to about 16 amino acids in length of the protein isoform of any of claims 4-6, that comprises non-exonic amino acid sequence.

8. A polynucleotide encoding the protein or fragment of any of claims 4-6 or the peptide of claim 7, or a fragment thereof at least 15, 20, 25 or 30 nucleotides in length, that optionally comprises non-exonic sequence, optionally wherein the polynucleotide is DNA or RNA.

9. A polynucleotide encoding the amino acid sequence of any of the amino acid sequences of Table 1A, or a fragment thereof at least 15, 20, 25 or 30 nucleotides in length, that optionally comprises non-exonic sequence, optionally wherein the polynucleotide is DNA or RNA, optionally the protein-coding sequence of any of the nucleotide sequences of Table 1A.

10. An expression vector comprising the polynucleotide of claim 8 or 9, or a fragment thereof at least 20 nucleotides in length, operably linked to one or more heterologous regulatory control nucleotide sequences, optionally wherein the heterologous regulatory control nucleotide sequence is a promoter, a transcriptional transactivator, an enhancer, a translation optimizing sequence, and / or a polyadenylation signal.

11. An oligonucleotide therapeutic agent that targets a polynucleotide comprising a nucleotide sequence encoding any of the amino acid sequences of Table 1A, optionally a nucleotide sequence of Table 1A, and which does not target the canonical gene transcript, optionally wherein the oligonucleotide therapeutic agent is antisense RNA, RNAi, shRNA, microRNA, saRNA, aptamer, ribozyme, or SSO.

12. The oligonucleotide therapeutic agent of claim 11 that is at least 10, 15, 20, 25 or 30 nucleotides in length and that comprises (i) a nucleotide sequence that binds to or is complementary to a fragment of any of the nucleotide sequences of Table 1A, orwherein said fragment optionally comprises a junction of exonic sequence and non- exonic sequence or about 40 to about 200 nucleotides of adjacent sequence, or (ii) a nucleotide sequence that is complementary to a region comprising a putative splicing site of any of the nucleotide sequences of Table 1A; or optionally (a) that targets any of the nucleotide sequences of PTEN JET-ORF, optionally for use in inducing an inflammatory response, or increasing IL-6 expression; or (b) that targets any of the nucleotide sequences of WWOX JET-ORF, optionally for use in the modulation of ATF6, KLF5, TWIST1 and TP73 transcriptional responses.

13. A host cell comprising the expression vector of claim 10.

14. A method of using the host cell of claim 13 to produce a protein isoform, or fragment thereof.

15. A population of dendritic cells or antigen presenting cells, optionally autologous, that have been pulsed with the peptide of claim 7, or that have been transfected with a polynucleotide of claim 8 or 9 or an expression vector of claim 10.

16. A vaccine or immunogenic composition capable of raising a specific immune cell response comprising:(a) a protein isoform or fragment of any of claims 4-6, or a peptide of claim 7, optionally with a physiologically acceptable buffer, carrier, or excipient, and / or optionally with an adjuvant or immunostimulant;(b) a polynucleotide of claim 8 or 9, or an expression vector of claim 10, or(c) a population of antigen presenting cells as defined in claim 15.

17. An antibody, or an antigen-binding fragment thereof, a T cell receptor (TCR), or a chimeric antigen receptor (CAR) that specifically binds a protein isoform or fragment of any of claims 4-6, or a peptide of claim 7, optionally in association with an MHC molecule, with a Kd affinity of about 10'6M or less, optionally wherein the antibody is a TCR-like antibody or the CAR is a TCR-like antibody -based CAR, optionally wherein the protein isoform comprises a transmembrane domain.

18. A method of producing an antibody, TCR or CAR that specifically binds a protein isoform or fragment of any of claims 4-6, or a peptide of claim 7, comprising the steps of:(a) contacting a population of antibodies, TCRs or CARs with said protein isoform or fragment or peptide, optionally a population produced by immunizing an animal with said protein isoform or fragment or peptide; and(b) selecting an antibody, TCR or CAR that binds to said protein isoform or fragment or peptide, optionally wherein said peptide is in association with an MHC or HLA molecule, or optionally wherein said protein isoform or fragment is expressed on the surface of a cell, wherein the antibody, TCR or CAR binds the protein isoform or fragment or peptide with a Kd binding affinity of about 10'6M or less.

19. An antibody, TCR or CAR produced by the method of claim 18.

20. An ex vivo method for producing a T cell comprising a TCR that specifically binds a protein isoform or fragment of any of claims 4-6, or a peptide of claim 7, useful in cell therapy, comprising:(a) stimulating T cells from a blood sample obtained from a patient with a composition comprising said protein isoform, fragment or peptide, or polynucleotide encoding said protein isoform, fragment or peptide, to prime, activate, and expand T-cells, optionally CD4+ T-cells, CD8+ T cells, or effector or central memory T cells,(b) optionally prior to step (a), the method comprises obtaining a sample of cells or tissue from the patient, optionally by leukapheresis or tumor biopsy, and confirming expression of said protein isoform or fragment thereof in said sample,(c) optionally subsequent to step (a), the method comprises confirming the specificity and functionality, optionally cytokine production activity and cytolytic activity, of the induced T-cells.

21. An antibody, TCR or CAR according to any of claims 17 or 19, wherein said antibody is a multispecific antibody that further targets at least one immune cell antigen, or wherein said TCR is a soluble fragment of a TCR fused to an antibody fragment that targets at least one immune cell antigen, optionally wherein the immune cell antigen is from a T cell, a NK cell or a dendritic cell, optionally wherein the immune cell antigen is CD3, CD16, CD30 or a TCR.

22. An immune cell comprising an antibody, TCR or CAR according to any one of claims17 or 19 or 21, optionally wherein the immune cell is defective for the Suv39hl gene.

23. The immune cell of claim 22, which is an allogenic or autologous cell selected from T cells, Natural Killer T cells, CD4+ T cells, CD8+ T cells, CD4+ / CD8+ T cells, TILs / tumor derived CD8 T cells, central memory CD8+ T cells, Treg, Mucosal- Associated Invariant T cells (MAIT), alpha / beta T cells, gamma / delta T cells, human embryonic stem cells, pluripotent stem cells, and / or myeloid or lymphoid lineage cells.

24. A pharmaceutical composition comprising an effective amount of the protein isoform or fragment of any of claims 4-6, or a peptide of claim 7, the polynucleotide of claim 8 or 9, the expression vector of claim 10, the oligonucleotide therapeutic agent of claim 11 or 12, the host cell of claim 13, the population of dendritic cells or antigen presenting cells of claim 15, the vaccine or immunogenic composition of claim 16, or the antibody, TCR or CAR of any of claims 17, 19 or 21, or the immune cell of claim 22 or 23, in combination with a sterile pharmaceutical excipient.

25. The protein isoform or fragment of any of claims 4-6, or a peptide of claim 7, the polynucleotide of claim 8 or 9, the expression vector of claim 10, the oligonucleotide therapeutic agent of claim 11 or 12, the vaccine or immunogenic composition of claim 16, or the antibody, TCR or CAR of any of claims 17, 19 or 21, for use in therapy, optionally for use in inhibiting cancer cell proliferation, or for use in cancer vaccination therapy of a subject, or for treating cancer in a subject, or optionally for use in treating a disease associated with the protein isoform sequence in Table 1 A, 3, or 4.

26. The host cell of claim 13, the population of dendritic cells or antigen presenting cells of claim 15, or the immune cell of claim 22 or 23, for use in cell therapy, optionally for use in cell therapy of cancer, or optionally for use in treating a disease associated with the protein isoform sequence in Table 1 A, 3, or 4.

27. The use of claim 25 or 26 which is for use in cancer in combination with at least one further therapeutic agent, optionally a chemotherapeutic agent, or an immunotherapeutic agent, optionally a checkpoint inhibitor.

28. The oligonucleotide therapeutic agent of claim 13 or 14 for use in modulating expression of and / or controlling function of the protein isoform of any of claims 4-6, optionally for inhibiting expression of the protein isoform, optionally for use in treating a disease associated with the protein isoform sequence in Table 1 A, 3, or 4.

29. The oligonucleotide therapeutic of any of claim 13 of 14 for use in modulating expression of and / or controlling function of the canonical protein associated with the protein isoform sequence of any of claims 4-6 in Table 1 A, 3, or 4.

30. An antibody, including an antigen-binding fragment thereof, for use in inhibiting the activity of the protein isoform of any of claims 4-6, optionally for use in treating a disease associated with the protein isoform sequence in Table 1 A, 3, or 4.

31. A method for determining the prognosis of a cancer or for detecting cancer cells characterized by increased expression of a protein isoform of Table 3 or 4, the method comprising the step of detecting or quantifying a protein isoform of any of the amino acid sequences of Table 1A, or a nucleic acid encoding any of the amino acid sequences of Table 1A, in a biological sample isolated from a patient, optionally wherein the detecting comprises:(I) (a) contacting said nucleic acid from the biological sample with a probe which specifically hybridizes to any of the nucleotide sequences of Table 1A, preferably a probe that hybridizes to the junction of exonic sequence and non-exonic sequence;(b) optionally detecting a complex formed between the probe and the nucleic acid from the biological sample, and comparing the amount of the complex in the biological sample to the amount of the complex in a reference sample comprising corresponding normal tissue;(c) wherein detecting said complex in the biological sample is associated with a better or worse prognosis of cancer,(d) or wherein detecting an increased amount of said complex in the biological sample as compared to the corresponding normal tissue indicates that the biological sample comprises cancer cells; or optionally wherein the detecting comprises(II) contacting the biological sample with an antibody that specifically binds the protein isoform of any of the amino acid sequences of Table 1 A, and detecting binding of the antibody with the protein isoform.

Citation Information

Patent Citations

  • Constitutive expression of costimulatory ligands on adoptively transferred T lymphocytes

    EP2537416A1

  • Machine foe cutting eur fedm skins

    US131A

  • Artificial antigen presenting cells and methods of use thereof

    US20020131960A1

  • Novel nucleic acids and polypeptides

    US20050196754A1

  • Method of controlling administration of cancer antigen

    US20130149337A1