Reporters constructs, systems and assays
A reporter system with a degradation-tagged fusion protein addresses the limitations of existing methods by enabling sensitive and flexible detection of antigen-specific T cells, facilitating the isolation and characterization of immunogenic peptides and T cell receptors for vaccine design and T-cell products.
Patent Information
- Application Number
- PCT/GB2025/051184
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-31
- Filing Date
- 2025-05-30
- Publication Date
- 2025-12-04
AI Technical Summary
Existing methods for detecting antigen-specific T cells, particularly CD4 T cells, are limited by the complexity of producing correctly folded MHC class II molecules and the instability of pMHC-T cell interactions, requiring individual production of peptide-MHC multimers for each peptide and being restricted to common MHC alleles.
A reporter system using a reporter fusion protein with a selection marker fused to a degradation tagging sequence, such as FKBP12F36V, which is degraded upon protease activity, allowing sensitive and flexible detection of T cell activation through protease recognition sequences like granzyme B or caspase-3/7, suitable for various proteases and applications like high-throughput drug screening.
The system provides highly sensitive and specific detection of antigen-specific T cells, enabling efficient isolation and characterization of immunogenic peptides and T cell receptors, suitable for vaccine design and T-cell product development.
Smart Images

Figure GB2025051184_04122025_PF_FP_ABST
Abstract
Description
[0001] REPORTERS CONSTRUCTS, SYSTEMS AND ASSAYS
[0002] The project leading to this application has received funding from the European Union’s Horizon 2020 research and innovation programme under grant agreement No 964998.
[0003] FIELD OF THE DISCLOSURE
[0004] The present disclosure relates to methods, systems, products and kits for in vitro detection of cellular events associated with protease activity, including but not limited to antigen recognition by T cells.
[0005] BACKGROUND
[0006] Recognition of peptides in the context of major histocompatibility complex (MHC) class I and II molecules by T cells is a very important component of cellular immunity against intracellular pathogens and cancer cells. This recognition is mediated through T cell receptors (TCR) on the surface of T cells, which are highly antigen-specific. Recognition of the pMHC (peptide-MHC complex) by the TCR activates the T cell, triggering a cascade of events including release of cytokines and other effector proteins such as granzyme B and perforins, that can ultimately induce apoptosis in cells displaying the recognised pMHC. The TCR-peptide MHC interaction has relatively low affinity and typically fast off-rate, making it difficult to detect antigen-specific T cells using single pMHC complexes. Therefore, MHC multimers are frequently used in this context, where specific pMHC ligands are multimerised to make soluble peptide-MHC multimers (e.g. tetramers, see Altman et al. 1996). Antigen-specific T cells that recognise these MHC multimers can be isolated by flow cytometry then characterised.
[0007] There are a number of limitations associated with MHC multimers. MHC multimers mostly include MHC class I molecules and as such are primarily suited for detection of CD8 T cells. This is because MHC class II molecules (for detection of CD4 T-cells) are more complex molecules to produce in a correctly folded state in vitro, and TCR affinities of CD4 T cells are generally lower than for CD8 T cells because the pMHC-T cell interaction is less stable. Further, all pMHC multimer technologies require the identification of the right peptide to reflect an antigen (i.e. the peptide that is likely to be produced by antigen processing and presented by MHC molecules on the surface of cells) - which is far from trivial for many class I MHC alleles and for class II MHC alleles, as well as the production of a correctly folded peptide loaded MHC multimer in vitro (i.e. for each peptide and MHC to be tested, it is necessary to separately generate a pMHC multimer). It is now possible to produce such complexes with greater ease for a variety of peptides by producing MHC complexes for a particular MHC allele loaded with a photolabile peptide which can be cleaved by UV light and exchanged for a peptide of choice. This still needs to be performed individually for each peptide, and is severely restricted in practice to common MHC alleles for which such UV- sensitive peptide-MHC complexes are available.
[0008] Thus, there is a need for improved methods to analyse antigen-specific T cells. SUMMARY
[0009] The present inventors hypothesised that a reporter of activation of T cells by antigens presented by antigen presenting cells could be developed which depends on the activity of granzyme B which is a specific marker of T cell activation. Reporter systems to identify T cells reacting with a target cell have been proposed to address some of the limitations of multimers. For example, Kula et al. (2019) and WO 2018 / 227091 described the T-Scan Report system in which T cells are co-cultured with target cells expressing a library of candidate antigens, where the target cells express a fluorogenic GzB reporter. The reporter comprises two domains of an IFP (infrared fluorescent protein) separated by a GzB cleavable constraining linker sequence, each IFP domain being also linked to a GFP domain such that in the absence of GzB, cells successfully transformed with the reporter construct are GFP positive but IFP negative, and in the presence of GzB these cells are GFP positive and IFP positive. Fluorescent cells are sorted by fluorescence activated cell sorting (FACS), allowing identification of antigens recognised by the T cells in the co-culture by PCR and sequencing of the antigen cassette of the isolated cells.
[0010] The present inventors have identified a number of limitations with this approach. Firstly, the approach has limited sensitivity due to low fluorescence intensity and the presence of background autofluorescence signal. Secondly, the approach is relatively inflexible since it requires the use of a fluorescent protein that can be separated into two domains with functionality restored only upon cleavage of a linker between the two domains, limiting severely the possibilities in terms of both the fluorescent protein and the linker sequence used.
[0011] The inventors therefore set out to design a new reporter system that is more sensitive, more flexible, and simpler to use. This was achieved by designing a reporter fusion protein in which a selection marker (here demonstrated as GFP) is fused in-frame to a degradation tagging sequence (here demonstrated as a FKBP12F36V degradation tagging sequence), with the two entities spaced with the granzyme B cleavage sequence. The inventors showed that cells transduced with this construct strongly express GFP. On the addition of the dTAG molecule, which is a PROTAC- based heterobifunctional degrader that binds the FKBP12F36V sequence, the protein is brought to the CRBN proteolysis system and rapidly degraded. The inventors have shown that K562 cells transduced with this construct strongly express GFP and become GFP negative after 4 hrs of dTAG-13 treatment, remaining so for at least 3 days. When these cells are pulsed with immunogenic CMV peptide and co-cultured with CMV specific T-cells, T-cells engage with K562 target cells, secrete Granzyme B, which cleaves the Granzyme B peptide sequence, releasing GFP from the FKBP12F36V tag so it is no longer degraded. Through proof-of-principle experiments the inventors showed the system to be highly sensitive and specific. Since different proteases recognise different peptide sequences, this method is also suitable to detect activity of different proteases of interest. For instance, a similar construct using a caspase-3 / 7 recognition sequence instead of a gzmb recognition sequence can be utilised as a reporter of apoptosis, useful e.g. for high throughput drug screening assays in cancer cells.
[0012] Thus, according to a first aspect, there is provided an isolated nucleic acid encoding a reporter fusion protein comprising a selection marker, a protease recognition sequence and a degron. Embodiments of the present aspect may have one or more of the following features.
[0013] The isolated nucleic acid may comprise a sequence encoding a selection marker in frame with a sequence encoding a protease recognition sequence and a sequence encoding a degron.
[0014] The degron may be a protein or peptide that leads to inducible degradation of the degron containing fusion protein. The degron may be a peptide or protein that targets the reporter fusion protein for proteasomal degradation in the presence of a degradation ligand. The degron may be a peptide or protein that binds to a degradation ligand, and the degradation ligand may be a molecule that binds to the degron and a ubiquitin ligase. The degron may be a mFKB12 peptide and the degradation ligand may be a dTAG molecule, optionally dTAG-13. The degron may be an IKZF2 peptide, an IKZF3 degron tag, or a CK1 a protein or peptide and the degradation ligand may be lenalidomide or a lenalidomide analogue. The degron may be an auxin-inducible degron and the degradation ligand may be an auxin or auxin analogue. The degron may be a Small Molecule- Assisted Shutoff (SMASh) tag and the degradation ligand may be an inhibitor of a protease included in the SMASh tag, optionally an inhibitor of NS3 protease. The degron may be a degron of a SMASh tag, optionally a NS4A peptide. The degron may be a proteolysis targeting chimera (PROTAC) target protein or peptide and the degradation ligand may be a PROTAC comprising a ligand for the target protein or peptide. The degron may be a HaloTag and the degradation ligand may be a HaloPROTAC.
[0015] The selection marker may be a positive selection marker. The selection marker may comprise a fluorescent marker. The selection marker may comprise a cell surface marker. The selection marker may comprise a cell survival protein. The fluorescent marker may be a fluorescent protein, optionally selected from GFP, mCherry and mStayGold. The fluorescent protein may be GFP. The cell surface marker may be a protein that is expressed on the surface of cells, optionally a CD34 protein, a CD19 protein, a NGFR protein, or a PD-L1 protein. The cell survival protein may be an antiapoptotic protein, optionally selected from BCL2, MCL1 and XIAP, an immune suppression protein, optionally PD-L1 , or a mutant protein that confers resistance to a cytotoxic compound or composition, optionally a mutant BCR-ABL protein.
[0016] The protease recognition sequence may be an amino acid sequence that is specifically recognised by a protease selected from granzyme B, gamma secretase, and a caspase, optionally caspase 3, 7, 8 and / or 10. In embodiments, the protease is granzyme B and the protease cleavage sequence is a sequence comprising amino acid sequence IEXD (SEQ ID NO: 23) where X can be any amino acid, VEXD (SEQ ID NO: 24) where X can be any amino acid, IGXD (SEQ ID NO: 25) where X can be any amino acid, or VGXD (SEQ ID NO: 26) where X can be any amino acid. In embodiments, the protease is caspase-3 / 7 and the protease cleavage sequence is a sequence comprising amino acid sequence DXXD (SEQ ID NO: 18) or DEXD (SEQ ID NO: 27) where X can be any amino acid. In embodiments, the protease is caspase-8 / 10 and the protease cleavage sequence is a sequence comprising amino acid sequence LQTD (SEQ ID NO: 4), LETD (SEQ ID NO: 29), VETD (SEQ ID NO: 30), LEHD (SEQ ID NO: 31), or VEHD (SEQ ID NO: 32). In embodiments, the protease is gamma secretase and the protease cleavage sequence is a sequence comprising amino acid sequence EDVGSNKGAIIGLMVGGVVIATVIVITLVMLKKK (SEQ ID NO: 33), FMYVAAAAFVLLFFVGCGVLL (SEQ ID NO: 34), LLYLLAVAWIILFIILLGVI (SEQ ID NO: 35), LPLLVAGAVLLLVILVLGVMV (SEQ ID NO: 36) or PVLCSPVAGVILLALGALLVL (SEQ ID NO: 37), or a homologue or variant thereof that is recognised by gamma secretase.
[0017] In embodiments, the protease recognition sequence is a peptide comprising a protease recognition motif, wherein the protease recognition motif is exposed in the reporter fusion protein. The protease recognition sequence may form an exposed loop comprising the sequence recognition motif. The protease recognition sequence may comprise a portion of the sequence of a native substrate of the protease, said portion comprising the protease recognition motif of the protease, or a sequence having at least 80%, 85%, 90%, or 95% sequence identity to said sequence provided that the protease recognition motif is maintained. In embodiments, the protease is granzyme B and the native substrate is a BID protein. In embodiments, the protease is caspase- 3 / 7 and the native substrate is a PARP1 protein. In embodiments, the protease recognition sequence comprises a portion of the sequence of human BID comprising amino acids Leu47 to Ser78, or a homologue thereof or sequence having at least 80%, 85%, 90%, or 95% sequence identity to said sequence. In embodiments, the protease recognition sequence comprises a portion of the sequence of human PARP1 comprising amino acids Val202 and Asp217, or a homologue thereof or sequence having at least 80%, 85%, 90%, or 95% sequence identity to said sequence.
[0018] Also described according to a second aspect is a vector comprising the isolated nucleic acid of any embodiment of the first aspect. The vector may have any one or more of the following optional features. The vector may be a viral vector, optionally a lentiviral vector. The vector may further comprise one or more promoter sequences, optionally selected from: a CMV promoter, SFFV promoter and a hPGK promoter. The vector may further comprise one or more antibiotic resistance genes and / or one or more candidate antigen sequences.
[0019] Also described according to a third aspect is a reporter fusion protein encoded by the isolated nucleic acid of any embodiment of the first aspect or expressed from the vector of any embodiment of the second aspect.
[0020] Also described according to a fourth aspect is a cell comprising the vector of any embodiment of the second aspect, the nucleic acid of any embodiment of the first aspect, or the reporter fusion protein of any embodiment of the third aspect. The cell may be an antigen presenting cell and / or a cancer cell. The cell may be a MHC monoallelic cell. The cell may be an immortalised cell, optionally wherein the immortalised cell is from a cancer cell line, optionally a leukaemia cell line, optionally a k562 cell line. The cell may be a HLA negative cell, optionally from a K562 cell line. The HLA negative cell may be genetically modified to express any HLA-allele of interest, permitting an easily customisable reporter to screen for antigen hits across multiple HLA types. This provides an off-the-shelf screening platform which could be used for example: in patient HLA-allele matched screens of immunogenic antigen identification for vaccine design, or antigen-specific T-cell products. Suitably, the HLA negative cell may be transduced with a vector encoding an HLA-allele of interest. The vector encoding the HLA-allele of interest may further encode the isolated nucleic acid of any embodiment of the first aspect or the vector encoding the HLA-allele of interest may be used with a separate vector according to any embodiment of the second aspect of the invention. The DNA sequences of HLA alleles are known in the art, for example at the IPD-IMGT / HLA Database (http: / / www.ebi.ac.uk / ipd / imgt / hla / ). Suitably, the HLA negative cell may be obtained as described herein and / or may be transduced with an HLA-allele of interest as described herein (see Example 6). The obtained mono-allelic HLA bearing cells can then be modified to express one or more antigens as described herein. The cell may be a Granzyme-B resistant cell line. The use of such a cell line aims to maximise the application of the invention by preserving cells expressing the reporter fusion protein and recognised by T-cells. This would likely permit the identification of more immunogenic hits by next generation sequencing. Furthermore, cells expressing the reporter fusion protein and recognised by T-cells could be sorted and expanded for downstream applications such as immunepeptidomics to identify the minimal epitope or used in T-cell expansion protocols to enrich for reactive T-cells.
[0021] The cell according to the fourth aspect may have been single cell clone based on expression of the selection marker (e.g. fluorescent protein). Any suitable method for single cell cloning may be used in the practice of the present invention. Suitably, cells comprising the vector of any embodiment of the second aspect, the nucleic acid of any embodiment of the first aspect, or the reporter fusion protein of any embodiment of the third aspect may have been single cell sorted and expanded, followed by screening for the largest fold change in expression of the selection marker following treatment with the degradation ligand. The method for single cell cloning described herein may be employed (see Example 4).
[0022] Suitably, the cell according to the fourth aspect may be a HLA negative cell and may have been single cell cloned based on expression of the selection marker as described herein.
[0023] Also described according to a fifth aspect is a library of cells according to the fourth aspect, wherein the cells in the library together display a plurality of antigens from a library of antigens. The plurality of antigens may comprise at least 10, at least 20, at least 50, at least 100, at least 1000, at least 5000 or at least 10000 different antigen peptides. The cells in the library may be cells that have been modified to express one or more antigens of said plurality of antigens. Suitably, the cells in the library may have been single cell cloned as described herein prior to being modified to express one or more antigens of said plurality of antigens. Suitably, the cells may be derived from a HLA negative cell line, optionally wherein the HLA negative cell line has been transduced with an HLA- allele of interest, as described herein.
[0024] Also described according to a sixth aspect is a method of making a reporter cell or reporter cell library, the method comprising obtaining a cell or population of cells and genetically modifying said cell(s) to express a reporter fusion protein according to the third aspect, optionally wherein the reporter cell or library has the features of any embodiment of the fourth or fifth aspects. The method may further comprise the step of single cell cloning the genetically modified cells based on expression of the selective marker (e.g. fluorescent protein). Single cell cloning may be performed as described herein. Suitably, the cells may be derived from a HLA negative cell line, optionally wherein the HLA negative cell line has been transduced with an HLA-allele of interest, as described herein.
[0025] Also described according to a seventh aspect is a kit comprising: (i) an isolated nucleic acid according to any embodiment of the first aspect, a vector according to any embodiment of the second aspect, or a cell or cell library according to any embodiment of the fourth or fifth aspects and (ii) a degradation ligand associated with the reporter fusion protein, wherein a degradation ligand associated with the reporter fusion protein is a compound that binds to the degron and causes the degradation of the reporter fusion protein.
[0026] Also described according to an eighth aspect is a kit comprising: (i) an isolated nucleic acid according to any embodiment of the first aspect or a vector according any embodiment of the second aspect; and (ii) one or more cells, optionally wherein the one or more cells have the features of any embodiment of the fourth or fifth aspects. The kit may further comprise: (ii) a degradation ligand associated with the reporter fusion protein, wherein a degradation ligand associated with the reporter fusion protein is a compound that binds to the degron and causes the degradation of the reporter fusion protein.
[0027] In kits according to the seventh or eighth aspects, the degradation ligand may be a molecule that binds to the degron and a ubiquitin ligase. The degron may be a mFKB12 peptide and the degradation ligand may be a dTAG molecule, optionally dTAG-13. The degron may be an IKZF2 peptide, an IKZF3 degron tag, or a CK1 a protein or peptide and the degradation ligand may be lenalidomide or a lenalidomide analogue. The degron may be an auxin-inducible degron and the degradation ligand may be an auxin or auxin analogue. The degron may be a Small Molecule- Assisted Shutoff (SMASh) tag and the degradation ligand may be an inhibitor of a protease included in the SMASh tag, optionally an inhibitor of NS3 protease. The degron may be a degron of a SMASh tag, optionally a NS4A peptide. The degron may be a proteolysis targeting chimera (PROTAC) target protein or peptide and the degradation ligand may be a PROTAC comprising a ligand for the target protein or peptide. The degron may be a HaloTag and the degradation ligand may be a HaloPROTAC.
[0028] Also described according to a ninth aspect is a method of isolating an immunogenic peptide and / or a T cell receptor (TOR) reactive to a peptide, the method comprising: (i) incubating a population of cells or a library of cells according to any embodiment of the third or fourth aspects in the presence of a population of T cells and optionally a degradation ligand associated with the reporter fusion protein, the population of cells comprising cells presenting one or more of a plurality of candidate peptides on their cell surface in complex with a MHC molecule, optionally wherein the MHC molecule is encoded by a selected HLA allele; and (ii) selecting one or more doublets of cells that comprise a T cell and an antigen presenting cell using a selection method associated with the selection marker comprised in the reporter fusion protein, thereby isolating an immunogenic peptide and a T cell expressing a TCR reactive to the peptide in the context of the selected HLA allele.
[0029] Methods according to the present aspect may have any one or more of the following optional features.
[0030] In embodiments, the method comprising obtaining sequence data associated with the one or more selected doublets. In embodiments, the method comprises analysing the sequence data to identify one or more of: a TCR expressed by the T cell in a doublet, a candidate peptide sequence expressed by an antigen presenting cell in a doublet, and an MHC molecule expressed by an antigen presenting cell in a doublet. The sequence data can be single cell RNA or DNA sequencing data. The step of analysing sequence data is typically computer implemented. Indeed, the typically size and complexity of next generation sequencing data sets is such that analysis is beyond the capability of the human mind (e.g. typical datasets comprise thousands of reads that must each be aligned to a reference sequence that is typically in the order of 3Gb for a human genome).
[0031] Methods according to the present aspect can be used in the context of screening a library of candidate antigens for antigens that are immunogenic, the method comprising: obtaining a population of antigen presenting cells comprising cells presenting a plurality of peptides corresponding to the candidate antigens in the library on their cell surface and expressing the reporter fusion protein; and isolating one or more immunogenic peptides of the plurality of peptides using a method according to the present aspect, thereby identifying candidate antigens of the library of candidate antigens that are immunogenic. The library of candidate antigens can be a library of expression vectors each including the coding sequence for a peptide of the plurality of peptides. The expression vectors can be viral vectors and the coding sequences can be included as minigenes. The expression vectors can include coding sequence for a peptide that is processed by an antigen presenting cell to obtain a peptide of the plurality of peptides. The methods may further comprise sequencing nucleic acid material isolated from one or more of the selected doublets, thereby identifying a candidate peptide associated with the antigen presenting cell in the doublet as an immunogenic peptide and / or a TCR expressed by the T cell in the doublet as a reactive TCR, optionally wherein the sequencing is single cell RNA sequencing.
[0032] The step selecting one or more doublets of cells that comprise a T cell and an antigen presenting cell may be performed using a cell sorting assay, optionally fluorescence activated cell sorting (FACS) or magnetic cell sorting (MACS). The step of incubating is performed for at most 24 hours, between 12 and 24 hours, about 16 hours, between 4 and 12 hours, or about 4 hours, optionally comprising or followed by incubation for at least 2 hours or at least 3 hours in the presence of the degradation ligand.
[0033] The cells or library of cells may have been peptide pulsed with the one or more of the plurality of candidate peptides. The cells or library of cells may have been modified to express the one or more of the plurality of candidate peptides. The method may comprise introducing an expression vector in cells of the population of antigen presenting cells prior to the incubating, the expression vector comprising a sequence encoding for the peptide, optionally wherein the expression vector is a lentiviral expression vector. Thus, in embodiments, the method comprises modifying the population of antigen presenting cells prior to the incubating by contacting the population of immortalised cells with a library of expression vectors in conditions suitable for introducing the expression vectors into the cells. In embodiments, the antigen presenting cells are from a cancer cell line, optionally a leukaemia cell line, optionally a k562 cell line. In embodiments, the population of T cells comprises CD8 T cells. In such embodiments, the antigen presenting cells may express a selected class I MHC molecule. In embodiments, the population of T cells comprises CD4 T cells. In such embodiments, the antigen presenting cells may express a selected class II MHC molecule. In embodiments, the population of T cells comprises naive T cells or T cells that have been preexpanded in the presence of CD3 / CD28 activators and IL2. In embodiments, the population of antigen presenting cells have been modified to include one or more expression vectors from a library of expression vectors, each expression vector of the library comprising a sequence encoding for a respective peptide and optionally a barcode sequence, wherein the population of immortalised cells comprise cells that express at least 10, at least 20, at least 50, at least 100, at least 1000, at least 5000 or at least 10000 different peptides.
[0034] In embodiments, the method comprises: obtaining a population of antigen presenting cells expressing a plurality of peptides and the reporter fusion protein, and selecting a subset of the population of antigen presenting cells prior to the incubating based on the presence of the marker following incubation with a T cell population, thereby enriching the population of immortalised cells for immunogenic peptides.
[0035] The methods according to the present disclosure also find use in the context of a method of isolating one or more T cells that are specific to respective one or more antigens, the method comprising: obtaining a population of T cells and a population of antigen presenting cells comprising cells presenting one or more peptides corresponding to the one or more antigens on their cell surface and expressing the reporter fusion protein; and isolating one or more T cell of the population of T cells using the method according to the present aspect, thereby isolating one or more T cells that are each specific to an antigen of the one or more antigens. The population of antigen presenting cells can comprise a plurality of subpopulations of cells each expressing an MHC molecule encoded by a different selected HLA allele. The at least one of the one or more peptides can be a peptide obtained by processing of a longer antigen peptide, by an antigen presenting cell presenting the peptide. Obtaining the population of antigen presenting cells can comprise transfecting the population of immortalised cells with one or more expression vectors each including the coding sequence for a peptide of the one or more peptides, or incubating the population of immortalised cells with peptides comprising the sequence of the one or more peptides.
[0036] The methods according to the present disclosure also find use in the context of a method of identifying an immunogenic triplet comprising a peptide, MHC molecule and TCR, the method comprising: isolating an immunogenic peptide and a T cell receptor (TCR) reactive to the peptide using a method as described herein, thereby identifying the peptide as the peptide of the immunogenic triplet, the selected HLA allele as the MHC molecule of the immunogenic triplet, and the TCR reactive to the peptide as the TCR of the immunogenic triplet.
[0037] In embodiments, the population of antigen presenting cells comprises at least 5% of cells presenting a peptide that is recognised by one or more T cells in the T cell population, the one or more T cells representing at least 5% of the T cell population. In embodiments, the method has a minimum detection limit of 5% of antigen presenting cells presenting a peptide recognised by one or more T cells in the population representing at least 5% of the T cell population.
[0038] In embodiments, the incubating comprises incubating the population of immortalised cells and the population of T cells using a 10:1 to 1 :10 ratio of cell numbers, optionally using about the same number of cells from the two populations (1 : 1 ratio) or a higher number of T cells compared to the number of immortalised cells. In embodiments, the T cells are unmodified and / or primary cells.
[0039] In embodiments, the antigen presenting cells are HLA-null cells that have been genetically modified to express the selected MHC allele. In embodiments, the antigen presenting cells are cells that have been genetically modified to express a single MHC allele, wherein the MHC allele is a class I allele or a class II allele. In embodiments, the antigen presenting cells are from a MHC monoallelic cell line.
[0040] Also described according to a tenth aspect is a method of screening for a set of conditions to identify conditions that affect antigen recognition, the method comprising: performing the method of any embodiment of the ninth aspect using a reference condition and a condition of the set of conditions to be screened; and comparing the number and / or identity of the selected doublets, wherein differences in the number and / or identity of the selected doublets between the reference condition and the test condition is indicative of the test condition affecting antigen recognition, optionally wherein the condition is an incubation condition and / or a condition that the antigen presenting cells and / or T cells have been subjected to prior to incubation, optionally wherein the condition is exposure to an active compound or composition.
[0041] Also described according to an eleventh aspect is a method of training an immunogenicity prediction algorithm, wherein an immunogenicity prediction algorithm comprises a machine learning model that takes as input an antigen sequence and optionally a TCR sequence and / or MHC molecule sequence or pseudosequence and produces as output an indication of whether the antigen is likely to be immunogenic, the method comprising: screening a library of candidate antigens for antigens that are immunogenic using a method according to the method of any embodiment of the ninth aspect, thereby obtaining training data comprising candidate antigens of the library of candidate antigens and associated labels indicating whether each candidate antigen was identified as immunogenic, immunogenic triplets each comprising a candidate peptide of a library of candidate peptides, an MHC molecule and a TCR, and / or pairs of immunogenic antigens and associated TCRs; and training the machine learning model using said training data.
[0042] Also described according to a twelfth aspect is a method of determining the effect of one or more test conditions on apoptosis of a target cell population or gamma secretase activity in a target population, the method comprising: (i) obtaining a target cell population comprising cells according to the third aspect, wherein the protease cleavage sequence is a caspase recognition sequence; (ii) culturing the target cells in the presence off the one or more test conditions, and optionally the degradation ligand associated with the reporter fusion protein; (iii) selecting or selectively quantifying cells of the target cell population that express the reporter fusion protein using a selection method associated with the selection marker comprised in the reporter fusion protein; and (iv) comparing the selected cells or quantities of selected cells between the one or more test conditions and / or between the one or more test conditions and corresponding reference values, thereby determining the effect of the respective test conditions on apoptosis of the target cells, optionally wherein the one or more test conditions are selected from: exposure to one or more compounds or compositions, co-culture with one or more cell populations, and exposure to one or more physico-chemical perturbations.
[0043] BRIEF DESCRIPTION OF THE FIGURES
[0044] Figure 1A illustrates schematically a general reporter construct of the disclosure.
[0045] Figure 1 B shows examples of a construct of A using a granzyme B cleavage sequence.
[0046] Figure 1C shows an example of a sequence used as a protease recognition sequence in examples of the disclosure using a granzyme B cleavage sequence. Figure 1 D shows a specific example of a construct of A using a casase-3 / 7 cleavage sequence
[0047] Figure 1 E shows an example of a sequence used as a protease recognition sequence in examples of the disclosure using a caspase-3 / 7 cleavage sequence (E).
[0048] Figure 2A illustrates schematically the interaction between a T cell and an antigen presenting cell (APC) presenting a peptide in the context of a MHC class I molecule, and how this interaction can be detected using an assay according to embodiments of the disclosure.
[0049] Figure 2B illustrates schematically a set up for screening a library of antigens and / or T cells / TOR according to embodiments of the disclosure.
[0050] Figure 2C shows a schematic example of a granzyme B reporter. Artificial antigen presenting cells, such as B-cells or K562 cells, are transduced to express a GFP-mFKBP12 fusion separated by a Granzyme B cleavage sequence such as RIEAD. On addition of the dTAG molecule, the GFP is rapidly degraded. Cells that have been engaged by a specific T-cell will be showered with perforin and Granzyme B which cleaves the FKBP12 degradation tag from GFP, leading to GFP stabilisation.
[0051] Figure 3 is a flowchart illustrating schematically a method of identifying an immunogenic peptide and / or a reactive T cell receptor according to embodiments of the disclosure.
[0052] Figure 4 is a flowchart illustrating schematically a method of providing an immunotherapy.
[0053] Figure 5 shows an embodiment of a system for identifying immunogenic peptides and / or reactive TCRs and / or for providing an immunotherapy.
[0054] Figure 6A,B show results of dTAG-13 activity on KA2_dTAG reporter cell lines. KA2_dTAG reporter cells were either untreated of treated with 100nM, 500nM or 1 uM dTAG-13 for 0.5, 1 ,2 or 4hrs at 37°C. KA2_dTAG cells were subsequently measured by flow cytometry for GFP expression. Fig. 6A shows Raw flow cytometry data for GFP expression. Fig. 6B shows summarised flow cytometry data for GFP expression.
[0055] Figure 7A illustrates the experimental design applied for a proof-of-concept of the use of a reporter construct of the disclosure to detect T cell recognition of a CMV peptide.
[0056] Figure 7B illustrates the experimental design applied for a proof-of-concept of the use of a reporter construct of the disclosure to detect T cell recognition of a CMV peptide expressed as a tile, with detection by sequencing.
[0057] Figure 8A-E show results of the experiments illustrated on Figure 7. Figs. A, B show results where CleavER cells (K562 cells expressing a reporter as described herein) were either unpulsed (no peptide) or pulsed with 100nM of 9mer CMV-peptide for 1 hr at 37°C and washed twice before incubation with CMV-specific T-cells for 4hrs or 16h at a 1 :1 E:S or 10:1 E:S ratio. As a negative control, CleavER cells were also incubated without T-cells. After incubation, cells were either untreated or treated with 1 uM dTAG-13 for 3hrs at 37°C and subsequently measured by flow cytometry for CD107a degranulation on T-cells and GFP positive reporter expression. Each condition was run in triplicate and reporter survival was measured using a 40sec stopping gate, % survival was determined as a % events in cultures compared to untreated CleavER cells. Figs. 8C, D show that reporter activity is dependent of antigen expression and antigen-specific T-cell recognition. CleavER cells were either peptide pulsed with NLV peptide or transfected with a tandem mini gene (TMG) encoding CMV 29mer amino acid sequence to measure endogenous processing and presentation of NLV T-cell epitope. (TMG_NLV). TMG_NLV was linked to CFP via a 2A sequence to measure TMG expression in KA2. TMG_NLV expressing cells were then spiked into KA2_dTAG reporters at 50%, 20% and 5% and co-cultured with T-cells spiked with CMV-specific T-cells at similar frequencies, for 6hrs or 16hrs. Cells were then treated with 2uM dTAG-13 for 3hrs and subsequently measured by flow cytometry for GFP, CD107a and activation marker CD137 expression. Fig. 6C shows results of experiments where CMV-specific T-cells and CleavER cells expressing tandem minigene (TMG) encoding CMV T-cell epitope linked with CFP were spiked into CD8-T-cells or CleavER cells respectively at stated frequencies and antigen reporter activity was assessed by FACS as above. T-cell reactivity was measured by CD107a and CD137 expression. Fig. 6D. shows results of FACS analysis of CleavER cell reporter activity with orwithout antigen. CleavER cells were gated on CD86 expression. Fig. 6E shows results of scRNA analysis (unique molecular identifiers (UMI) counts for the CMV tile) of GFP positive and GFP negative CleavER cells spiked with 20% CMV antigen expressing CleavER cells following coculture with CMV-specific T-cells. CMV antigen was enriched on GFP+ CleavER cells.
[0058] Figure 9A-C shows results of the experiments illustrated on Figure 7A. Fig. 9A shows results of experiments where K562 cells were stably transduced with a GFP-GBC-FKBP12F36V construct where the GFP is driven from the CMV promoter in the PLX302 lentiviral backbone (GBC= granzyme B cleavage sequence RIEAD) and then pulsed with CMV peptide. Cells were co-cultured for 4 hrs with donorT-cells from a CMV+ donor at 1 :1 effector to sensor E:S ratio, then treated with 1 uM dTAG-13 for 3hrs. Flow cytometry plots for GFP vs side scatter (SS) in the various conditions are shown. Fig.9B is a bar graph showing %GFP+ cells from the experiment performed in biological triplicate. Fig. 9C shows results on T-cell activation as measured by CD107a expression by flow cytometry.
[0059] Figure 10A-C shows results of the experiments illustrated on Figure 7B without the sequencing step (i.e. cell sorting based results). Figs. 10A-B show that the use of a tile-based antigen expression system requires a longer co-incubation with T cells in order to mount a T cell response, but after 16 hours of co-incubation, a strong dose dependent T cell activation signal is detected as evidenced by the expression of activation markers CD107a and CD137 measured by flow cytometry. Fig. 10C shows %GFP positive doublets detected using the reporter as a function of the percent of target cells expressing the CMV antigen and the percentage of T cells that are reactive to the CMV antigen in the population.
[0060] Figure 11A shows an example vector of the disclosure.
[0061] Figure 11 B shows an example reporter fusion protein sequence of the disclosure.
[0062] Figure 12 shows a scatter plot showing bulk RNA analysis of POC CMV model. 2% CMV TILE+ CleavER cells were spiked into PM7+CleavER cells and co-cultured overnight with 2% CMV- specific T-cells. Following DTAG-13 treatment GFP+ and GFP- CleavER cells were FACS sorted, expanded for 7 days and bulk RNA sequenced to identify TILE expression using a DECOD bioinformatic pipeline. wqLog fold change between the tile counts in the GFPpos and GFPneg cell populations is plotted against the log of the mean read counts across the experiments. Log fold change values are corrected by the counts in No-T-cell control conditions to account for technical noise in the assay and to determine a positive hit calling threshold (orange line).
[0063] Figure 13 shows a scRNA downstream empirical bayes analysis workflow. Cleaver screen of 20k Neoantigen library (PM7) in HLA-A*02:01 across three healthy donors identified by scRNA seq. The percentage of cells in the GFPpos and GFPneg fractions are plotted against each other. Tiles are coloured by the normalised EB estimate (z-score) values assigned to each tile. Tiles with higher confidence of enrichment are shown in red; tiles with lower confidence of enrichment are shown in yellow. Tiles with a z-score value of 1 .5 or more was labelled a significant enrichment.
[0064] Figure 14 shows an upset plot outlining the number of tiles identified in each CleavER screen of 20k Neoantigen library (PM7) in HLA-A*02:01 across three healthy donors identified by scRNA seq. The total number of hits in each experiment are shown in the Total Hits side bar char. The number of tiles which intersect between H011 , H013 & H018 are shown and overlapping tiles are shown in the pink donor matrix below.
[0065] Figure 15 shows KA2_C LEAVE R_GFP single cell clone (SCC) out-performs bulk populations post DTAG-13 treatment. K562 expression HLA-A*02:01 , CD80 and CD86 were lentivirally transduced with GFP-CleavER cassettes (KA2_CLEAVER). HLA-A2hi, CD80hi, CD86hi and GFPhi expressing KA2_CleavER were then single cell sorted into 96 well plates and expanded for 2 weeks. A) Expanded SCC were then screened by FACS for the largest fold change in GFP expression following a 4hr DTAG-13 treatment. B) KA2_CLEAVER cells (Bulk and SCC B5) were treated with DTAG-13 for increased timepoints before analysis by FACS for GFP expression. C) FACS plots of GFP expression with and without 24hrs DTAG13 treatment.
[0066] Figure 16 shows single cell clone screening for mCherry and StayGold CleavER cells. A) mcherry_CleavER and B)Staygold_CleavER constructs were transduced into HLA A*02:01 expressing K562 (KA2_CLEAVER) and single cell cloned (SCC) based on fluorescent protein expression. SCC were screened for the largest fold difference in fluorescent protein expression when cells were treated for 4hrs with DTAG-13 treatment. Chosen clones were then treated with prolonged DTAG-13 treatment times to determine degradation kinetics.
[0067] Figure 17 shows suitability of an RQR8 cell surface reporter. A) RQR8_CleavER construct (SEQ ID:RQR8_gzb_Mfkb12) used to transduce HLA-A*02:01 + K562 cells. B) CD34 expression on RQR8_CleavER cells after various treatment times with DTAG-13 and representative FACS plots.
[0068] Figure 18 shows suitability of an NGFR8_CLEAVER cell surface reporter. NGFR_CleavER construct including a 23 amino acid intracellular domain of CD8a (NGFR8_GZB_MFKB12) transduced into HLA:A*02:01 + K562 cells. B) NGFR expression in bulk sorted NGFR_CleavER cells without DTAG-13 treatment and after 72hrs DTAG-13 treatment. C) Frequency (left) and GeoMean (right) of NGFR expression after prolonged treatment times (hrs) with DTAG-13. D) Fold difference in NGFR signal after 72hr DTAG-13 treatment of NGFR8_CleavER single cell clones (SCC). E) FACS plot of representative SCC C12 after 72hrs treatment, values represent the frequency of NGFR expressing cells.
[0069] Figure 19 shows suitability of auxin-mediated protein degradation as an alternative to DTAG-13 in CleavER cells. A) 2-protein Auxin-mediated protein degradation construct (SEQ ID: GFP_GZB_Maid) transduced into HLA-A*02:01 + K562 cells (KA2_CLEAVER_AUXIN). B) GFP expression in bulk sorted KA2_C LEAVE R_AUXIN vs KA2_CLEAVER_GFP_B5 (Mfkb12 DEGRON) following treatment with 500uM indole-3-acetic acid (IAA; a natural auxin). C) Frequency and D)Geomean of GFP expression in bulk sorted KA2_CLEAVER_AUXIN vs KA2_CLEAVER_GFP_B5 at increasing doses (uM) and increasing treatment time (hrs) with IAA.
[0070] DETAILED DESCRIPTION
[0071] Intracellular proteases are enzymes that cleave proteins, leading to their inactivation or activation. Two examples of particular interest are the granzyme and caspase proteases involved in cellular apoptosis. When a T-cell becomes activated after interacting with an immunogenic antigen presented on MHC class I, it secretes perforin, leading to the generation of cell membrane pores through which granzyme B can enter and kill the target cell. Granzyme B is a powerful protease that recognises and cleaves proteins that contain specific peptide sequences. Important protein substrates containing these sequences include BID and Caspase-3 respectively, cleavage of which leads rapidly to mitochondrial apoptosis and cell death. The ability to read out which cell is targeted by a particular T-cell is of interest to many in the field of immunology, virology, autoimmune disease and cancer. For instance, DNA mutations in cancers are frequently expressed through MHC as mutant peptides on the cell surface as neoantigens. From a therapeutic point of view, identifying the exact T-cell receptor (TCR) alpha-beta sequence that recognises a specific neoantigen enables the generation of cellular therapies such as tumour infiltrating lymphocytes (TILs), cancer vaccines and chimeric antigen receptor (CAR) T-cells. However, the major challenge in the field has been matching which specific TCR alpha-beta pairs react with which neoantigen, with only a few fully validated interactions documented. Some progress has been made by expressing massive libraries of potential neoantigens in cells in vitro and co-culturing these cells with polyclonal T-cells. However, there is still a major technical challenge in capturing activated T-cells in contact with their target cell, particularly as there tend to be lots of non-specific ‘sticky’ interactions between cells in these assays.
[0072] To overcome this challenge, the present inventors have developed a reporter system where a target cell becomes selectable when it is specifically targeted by a reactive T-cell. A general embodiment is illustrated on Figure 1A. In embodiments, marker expression cells (e.g. fluorescent cells) can then be isolated by cell sorting (e.g. fluorescence activated cell sorting, FACS) as doublets (target cell and its specifically interacting T-cell) and the antigen and TCR alpha-beta pairs identified by single cell sequencing. An embodiment of the reporter system includes a GFP chimaera where the GFP sequence is fused in-frame to the FKBP12F36V degradation tagging sequence, with the two entities spaced with the granzyme B cleavage sequence (Figure 1B, top). Cells transduced with this construct strongly express GFP. On the addition of the dTAG molecule, which is a PROTAC-based heterobifunctional degrader that binds the FKBP12F36V sequence, the protein is brought to the CRBN proteolysis system and rapidly degraded (Figure 2C). When transduced into antigen presenting cells co-cultured with T-cells, upon specific recognition of the antigen(s) presented by the cells (Figure 2A), the T cells engage with the antigen presenting target cells, secrete Granzyme B, which cleaves the Granzyme B peptide sequence, releasing GFP from the FKBP12F36V tag so it is no longer degraded (Figures 2A, 2B, 2C). The approach has very good sensitivity and is simple to implement, making it amenable to screening of antigens for immunogenicity at scale, for example using a set up as illustrated on Figure 2B. Further, since different proteases recognise different peptide sequences, the approach can be adapted to detect activity of different proteases of interest. For instance, a construct encoding a selectable marker, caspase-3 / 7 recognition sequence and degron (e.g. GFP-DEVD- FKBP12F36V) could be utilised as a reporter of apoptosis, useful for high throughput drug screening assays.
[0073] The methods described herein have many advantages. For example, the methods described herein can be used in large-scale neoantigen screening, and large-scale drug screening assays, because of their simplicity of implementation and high sensitivity. The former in particular is currently an extremely challenging task because of the extremely high diversity of T cell repertoires and antigens, leading to a combinatorial explosion when screening libraries of antigens against diverse T cell populations. This makes the task of identifying those rare antigen-T cell engagement events extremely difficult, especially in the presence of noise, background signal and non-specific interactions. The present approach is able to overcome these challenges by ensuring extremely high sensitivity and specificity of detection of antigen presenting cells that have been recognised by an activated T cell. Additionally, the approach is simple to use. Markers such as fluorescent protein constructs can be easily cloned and expressed in diverse cell types, and the present reporter system uses a single vector encoding a single fusion protein. Additionally, because the present approach has no conformational requirement other than for the protease recognition sequence to be cleavable, it is very flexible in terms of the sequence connecting the marker and degron, and in terms of the 3D conformation of this sequence in different intracellular environments. By contrast, approaches such as those in Kula et al. (2019) and WO 2018 / 227091 require a specific configuration of the linker between the two IFP domains. Further, the approach is highly versatile. The approach is usable as a reporter for any protease for which a cleavage sequence peptide is known, with any selection marker for which a selection technology exists, including fluorescent markers (which can be used to selected cells by e.g. FACS) and cell surface expressed markers that lead to cell surface expression of an epitope (e.g. CD34, PD-L1 ) that can be used for FACS or any pull down technology such as magnetic cell sorting, or survival markers (e.g. expression of a survival protein such as BCL2, PD-L1 or MCL1). Multiple selection strategies can even be combined, such as e.g. using a fusion protein comprising a fluorescent or cell surface marker in cells that also express a survival protein, or using a cell surface expression marker that is also a survival marker (e.g. PD-L1). Further, the methods described herein do not rely on relocalisation of a marker upon exposure to the protease. Instead, the marker is either present (in the presence of the protease) or absent (in the absence of the protease). This is by contrast with other known gzb reporters such as that in Liesche et al. (2018), which relies on a nuclear export signal excluding a fluorescent marker from the nucleus in the absence of gzb. Sorting positive cells using the methods described herein can be done in a high throughput, non-destructive manner, enabling recovery and identification of the target cells or parts thereof (e.g. antigens expressed by positive target cells), which is not possible using localisation based assays (requiring confocal microscopy). Additionally, the methods described herein are compatible with single cell readouts, such as e.g. single cell sequencing. Many available readouts for cytotoxicity (e.g. luciferase based caspase-3 / 7 assays) provide readouts at the level of cell culture wells. By contrast, individual cells can be selected using the reporters described herein and individually analysed, enabling more precise quantification of effects, and characterisation of the positive cells or associated cells (e.g. T cells forming a doublet with a positive antigen presenting cell). Further, specific identification and selection of positive cells can be performed using the reporters of the present disclosure using single selection mechanisms, for example using a single fluorescent channel based identification (when the reporter comprises a fluorescent marker). This is by contrast to e.g. the approach in Kula et al. (2019) which requires sorting on both GFP and IFP signals.
[0074] Embodiments that use the dTAG system additionally benefit from experimentally demonstrated high robustness of degradation of the reporter fusion protein, leaving very little background marker signal, thereby enabling highly sensitive and specific detection. Such embodiments additionally benefit from the dTAG compound being recycled such that the compound does not need to be) repeatedly added to the cell culture. Such embodiments additionally benefit from rapid degradation dynamics, enabling use of the system in short term co-culture assays. Note that the use of the dTAG system is not a requirement for this and any system that enables rapid degradation of a degron containing protein would be usable. Rapid degradation dynamics are particularly beneficial in the context of antigen-T cell recognition or any other context in which cytotoxicity I apoptosis is expected (and indeed a consequence of the cellular events that are being detected), since targeted cells can otherwise be lost from the system prior to detection.
[0075] The present disclosure provides novel assays for screening of antigen recognition events. It was developed in the context of a project in which it was desirable to screen very large libraries of antigens against primary T cell populations to identify the hallmarks of an immunogenic tumour neoantigen (iNeoAg) as well as enable the development of predictors of immunogenicity and immunogenic triplets. The present inventors found that none of the existing platforms for performing immunogenicity screening enabled them to do this. They therefore set out to design a new screening platform that could identify T cells that recognise an antigen, and / or antigens that are recognised by T cells. In this context, the use of the reporter system described herein, as summarised above and further demonstrated below, has a number of additional advantages compared to the traditional approach to identify T cells that recognise an antigen, i.e. MHC multimers. Indeed, the reporter can be used to screen for antigen recognition events in the context of a very wide range of HLA allele (essentially any that can be cloned), when used in the context of a monoallelic antigen presenting cell. Further, the reporter can be used to screen antigen libraries in a high throughput manner. By contrast, MHC multimers are only commercially available for a handful of common HLA alleles (approximately 30, which is far smaller than the complete human HLA diversity - 38,416 HLA and related alleles described by the HLA nomenclature and included in the IPD-IMGT / HLA Database), because the process of expressing and correctly folding pMHC complexes in vitro is complex, time consuming and costly. For the same reason, MHC multimers are typically limited to class I alleles. By contrast, the present assay is amenable to both class I and class II allele screening, enabling the identification of both CD8+ and CD4+ (cytotoxic CD4+) T cell reactivities. Additionally, the approach enables to sensitively and specifically measure T cell activation, rather than just binding. For a long time it was assumed that the two events were equivalent but emerging research has shown that binding does not necessarily lead to activation in real (e.g. patient-derived) T cell populations. This is particularly important at least in the context of personalised medicine that relies on T cell activation (including e.g. many cancer vaccines and T cell therapy approaches).
[0076] Thus, the methods and products of the present disclosure find uses in the field of immune- oncology, for example for the identification of neoantigens for neoantigen-based therapies, including cell-based therapies and vaccine based therapies, for the identification, development and / or characterisation of T cell receptor based therapies (e.g. CAR T- cell therapies, bispecific T cell engagers, engineered T cell therapies, etc.), and for the identification, development and / or characterisation of compounds or compositions that effect a T cell mediated immune response against cancer cells.
[0077] Further, the methods and products of the present disclosure also find uses in the field of immunology and infectious disease biology, for example for the identification of antigens for vaccination, or identification, development or characterisation of compounds that effect a T cell mediated immune response against a pathogen. Further, the methods and products of the present disclosure also find uses in the field of autoimmunity, for example for the identification of antigens that mediate an autoimmune response, and for the identification, development and / or characterisation of compounds or compositions that effect a cell mediated autoimmune reaction.
[0078] Further still, the methods and products of the present disclosure also find uses in the field of oncology, for example for the identification, development and / or characterisation of compounds or compositions that induce apoptosis in target cells, including e.g. specific cell types of interest such as cancer stem cells.
[0079] In the present disclosure, the following terms will be employed, and are intended to be defined as indicated below.
[0080] Figure 1A illustrates schematically a general reporter construct of the disclosure. The reporter includes a selection marker 10, a protease recognition sequence 12, and a degron 14. The term “reporter construct” (also referred to simply as “reporter”) refers both to a fusion protein comprising a selection marker 10, a protease recognition sequence 12, and a degron 14, and to a nucleic acid comprising a sequence encoding such a fusion protein.
[0081] A selection marker refers to a protein that enables isolation (i.e. selection) of cells that express the protein. The selection marker is preferably a positive selection marker. A positive selection marker is a marker that enables positive selection of cells that express the marker. Any protein that when present in cells enables the cells to be enriched compared to cells in which the protein is not present can be used in this context. Positive selection (enabling enrichment of positive cells) has a wider dynamic range than negative selection (because enrichment of a signal associated with positive cells is more reliably detected than depletion), and is easily multiplexed. A selection marker may be a fluorescent marker, such as a fluorescent protein. Any fluorescent protein known in the art may be used, such as e.g. GFP, StayGold (Hirano et al. 2022), CFP, YFP, etc. The use of StayGold (also referred to as mStayGold) is advantageous because the protein has a brightness of 136 (compared to 33.5 for GFP), leading to a stronger shift in fluorescence upon cells becoming positive for the selection marker. A selection marker may be a cell surface marker. A cell surface marker refers to a protein that is expressed on the cell surface. A cell surface marker may be any protein that can be expressed on the surface of cells, including any naturally occurring cell surface protein and any engineered protein comprising a cell surface localisation signal sequence and / or a transmembrane domain (e.g. a type I transmembrane protein). A cell surface marker may be a CD34 protein, a CD19 protein, or a nerve growth factor receptor (NGFR, also known as tumor necrosis factor receptor superfamily member 16, LNGFR and CD271). A CD34 protein refers to a protein that comprises at least part of the sequence of CD34 (optionally a mammalian, such as human CD34 or mouse CD34, e.g. Uniprot ID P28906 or Q64314, respectively for the human and mouse sequences) comprising a transmembrane domain, and an extracellular epitope. A CD34 protein may be a truncated CD34 protein (e.g. a CD34 splice variant as described in Fehse et al. 2000). A CD34 protein may be protein comprising a transmembrane domain and a minimal epitope for recognition by an anti-CD34 antibody. For example, the CD34 protein may be the protein RQR8 (a 136-amino-acid protein described in Philip et al. 2014, that is recognized by the anti-CD34 antibody QBEndl O). A CD19 protein (also known as B-lymphocyte surface antigen B4 and T-cell surface antigen Leu-12) refers to a protein that comprises at least part of the sequence of CD19 (optionally a mammalian, such as human CD19 or mouse CD19, e.g. Uniprot ID P15391 or P25918, respectively forthe human and mouse sequences) comprising a transmembrane domain, and an extracellular epitope. A CD19 protein may be a truncated CD19 protein, such as ACD19 described in Di Stasi et al. 2011 . A NGFR protein refers to a protein that comprises at least part of the sequence of NGFR (optionally a mammalian, such as human NGFR or mouse NGFR, e.g. Uniprot ID P08138 or Q9Z0W1 , respectively for the human and mouse sequences) comprising a transmembrane domain, and an extracellular epitope. A NGFR protein may be a truncated NGFR protein, such as ALNGFR as described in Rudoll et al. 1996.
[0082] A selection marker may be a cell survival protein. A cell survival protein refers to a protein that confers a survival advantage to cells expressing the protein compared to cells that do not express the protein. The survival advantage may be context dependent. For example, the survival advantage may be associated with exposure of the cells to a cytotoxic compound or composition that the cells would otherwise be sensitive to. As another example, the survival advantage may be associated with exposure to cytotoxic immune cells. The cell survival protein may be selected from: an antiapoptotic protein (e.g. BCL2 family protein, such as BCL2, MCL1 or XIAP), an immune suppression protein (e.g. PD-L1), or a protein that confers resistance to a cytotoxic compound or composition that the target cell is otherwise sensitive to. An immune suppression protein is a protein that, when expressed by a target cell (e.g. an antigen presenting cell) suppresses, down- regulates or inhibits an immune response. For example, the protein may inhibit T-cell activation. An immune suppression protein may be a ligand of an inhibitory immune checkpoint protein. An immune suppression protein can be selected from: PD-L1 (also known as CD274), PD-L2 (also known as PDCD1 LG2), CD80, CD86, TNFRSF14, PVR, Nectin-2, TIMD4, CEACAM1 , Galectin-9 (LGALS9), SLAMF1 , CD48, SLAMF6, SLAMF7, CLEC7A, CLEC10A, CD276, complement C1 q (C1 QA, C1 QB and / or C1 QC), VSIR, BTN2A2, BTN3A1 , BTNL2, SIGLEC15 and VTCN1 . All of these proteins have been documented to be expressed by antigen presenting cells and have a co- inhibitory effect for T-cell activation. In embodiments, an immune suppression protein is selected from: PD-L1 (also known as CD274), PD-L2 (also known as PDCD1 LG2), CEACAM1 , Galectin-9 (LGALS9), CLEC7A, CLEC10A, complement C1 q, VSIR, BTN2A2, BTN3A1 , BTNL2, SIGLEC15 and VTCN1. All of these proteins have been documented to be expressed by antigen presenting cells and have an exclusively co-inhibitory effect for T-cell activation. In embodiments, an immune suppression protein is selected from: PD-L1 (also known as CD274), CLEC7A, complement C1 q, VSIR, BTN3A1 , BTNL2, and VTCN1. In embodiments, the immune suppression protein is PD-L1 . In embodiments, the antiapoptotic protein or immune suppression protein is a protein that is not expressed or not constitutively expressed by the target cells. The use of an antiapoptotic or immune suppression protein advantageously ensures that positive cells do not undergo apoptosis (e.g. immune-mediated apoptosis such as e.g. following a T-cell recognition event). This in turns results in a positive selection of the cells under pro-apoptotic conditions, and also ensures that there is no caspase-activated DNase (ICAD) degradation of the DNA in the positive cells. This in turns improves the recovery of the positive cells and any relevant genetic material therein, without the need to introduce any additional constructs (such as e.g. a construct encoding a caspase resistant version of ICAD as was done in Kula et al. 2019. In embodiments, the cell survival protein is a drug resistance mutant protein. Any protein known in the art that confers resistance to a cytotoxic compound or composition, that is not normally expressed by the target cells in which the reporter is used can be used for this purpose. In embodiments, the cell survival protein is a mutant BCR-ABL (also referred to as BCR-ABL1 ) and the target cell has a BCR-ABL fusion that makes it sensitive to a cytotoxic drug (e.g. Imatinib) in the absence of the mutant BCR-ABL. For example, a mutant BCR-ABL protein may be BCR-ABL T315I. The presence of the BCR-ABL T315I mutant confers resistance to Imatinib or Dasatinib, such that any target cell that is sensitive to imatinib / dasatinib can be used in combination with reporters where the selection marker is BCR-ABL T315I. In such embodiments, the target cells carrying the reporter may be cultured in the presence of the cytotoxic compound or composition. The cytotoxic compound or composition can be added to the cell culture at the same time as a degradation ligand and / or at the same time as exposing the target cells to conditions expected to trigger the activity of the proteases to be detected. In embodiments, the cytotoxic compound or composition is added to the cell culture at the same time as exposing the target cells to conditions expected to trigger the activity of the proteases to be detected, or after said exposing. In embodiments, the reporter comprises a plurality of selection markers. In embodiments, the reporter comprises a selection marker that is or comprises a membrane anchored protein. Fluorescent or cell surface markers may be used for detection of positive cells using flow cytometry and / or for selection of positive cells using fluorescence activated cell sorting. In this context, positive cells or doublets of cells may be identified as those satisfying predetermined gating parameters that apply to the signal obtained from the single cell detection technology. Gating parameters may for example comprise the intensity of a signal associated with a fluorescent marker (either part of the reporter fusion protein or associated with an antibody that specifically binds to a cell surface marker that is part of the reporter fusion protein) being above a predetermined threshold. A “degron” (also referred to herein as “degradation tag”) refers to a protein or peptide that is recognised by a degradation mechanism and leads to degradation of the protein comprising the degron. Thus, the reporters described herein may also be referred to as degron-fused proteins or degron containing fusion proteins. The degradation of the protein comprising the degron may be constitutive or may be conditional (also referred to as inducible). For example, the degradation of the protein comprising the degron may be dependent on the presence of a specific molecule (large or small molecule) in the cell. In other words, the degron may be recognised by a degradation mechanism in the presence of an additional component (also referred to as degradation ligand or simply ligand), which can be a small molecule or a large molecule (e.g. another protein). The use of a conditional degradation mechanism advantageously enables to verify that the reporter is present in the cells, or even to select cells in which the reporter is present prior to using the cells. Further, the use of a conditional (inducible) degradation mechanism enables control of the degradation of the degron-fused proteins. The ligand typically is a molecule that recruits directly or indirectly components of the ubiquitin-proteasome pathway. In embodiments, the degradation of the protein comprising the degron is conditional on the presence of a compound that can be added to the cell culture medium in which the cells expression the reporter fusion protein are cultured. This advantageously enables faster dynamics and more precise control than relying on systems that require expression of a separate additional compound for degradation. In embodiments, the degron is a protein or peptide that binds to the degradation ligand, and the degradation ligand is a molecule that binds to the degron and to a ubiquitin ligase. In embodiments, the degron is a mFKB12 peptide, also referred to as mutant FKB12 peptide or FKBP12F36Vpeptide. In such embodiments, the degradation ligand can be a dTAG molecule. A dTAG molecule can be dTAG- 13, dTAG-7, dTAG-48, dTAG-51 or a dTAG analogue. Examples of dTAG molecules are provided in Nabet et al. 2018 and WO 2017 / 024319. The dTAG molecule is a heterobifunctional molecule that is able to bind to FKBP12F36V mutant sequence (mFKB12 peptide) with one arm, and the E3 ligase CRBN with the other, leading to rapid degradation of the target protein of interest. An exemplary embodiment is shown on Fig. 1B (top), which shows a specific embodiment in which the protease recognition sequence 12 is a granzyme B cleavage sequence, the selection marker 10 is a GFP protein, the degron 14 is a mFKB12 peptide, and the degradation ligand 16 is a dTAG molecule. The mFKB12 peptide interacts with the dTAG molecule leading to degradation of the fusion protein comprising the mFKB12 peptide by the proteasome. In embodiments, the degron is an IKZF2, IKZF3 degron tag (e.g. as described in Jan et al. 2021), or a CK1 a protein or peptide. In such embodiments, the additional degradation ligand can be lenalidomide or a lenalidomide analogue. Lenalidomide (Miyamoto et al. 2023; also known as DEG-77) is a bifunctional molecule based on a naphthamide scaffold that binds to target proteins IKZF2 and CK1a and the E3 ligase substrate adapter cereblon, leading to degradation of the target proteins. An exemplary embodiment is shown on Fig. 1 B (middle), which shows a specific embodiment in which the protease recognition sequence 12 is a granzyme B cleavage sequence, the selection marker 10 is a StayGold protein, the degron 14 is a IKZF2 peptide, and the degradation ligand 16 is a lenolidomide molecule. In embodiments, the degron is an auxin-inducible degron (AID). In such embodiments, the degradation ligand can be an auxin molecule. An auxin molecule can be an auxin (e.g. indole-3-acetic acid (IAA)) or auxin analogue (e.g. 5-phenyl-indole-3-acetic acid (5-Ph- IAA)). Auxin promotes the interaction between the AID and the TIR1-containing Skp1 -F-box-Cullin E3 ubiquitin ligase, ultimately leading to polyubiquitination and degradation of the AID-tagged protein. Auxin inducible degrons are described in Nishimura et al. 2009, Kubota et al. 2013 and Yesbolatova et al. 2020. An auxin inducible degron can be selected from AID described in Nishimura et al. 2009, mini-AID described in Kubota et al. 2013, and AID2 described in Yesbolatova et al. 2020. An illustrative auxin inducible degron is shown in SEQ ID NO: 54. Suitably, the illustrative auxin inducible degron may comprise SEQ ID NO: 54 or a variant with at least 80%, at least 85, at least 90%, at least 95% or at least 99% sequence identity thereto. In embodiments, the auxin inducible degron is AID2 (an OsTIR1 (F74G) mutant) and the degradation ligand is 5- phenyl-indole-3-acetic acid (5-Ph-IAA). An exemplary embodiment is shown on Fig. 1B (bottom), which shows a specific embodiment in which the protease recognition sequence 12 is a granzyme B cleavage sequence, the selection marker 10 is a StayGold protein, the degron 14 is an AID peptide, and the degradation ligand 16 is an auxin molecule. In embodiments, the degron is a Small Molecule-Assisted Shutoff (SMASh) tag or a degron of a SMASh tag. Examples of SMASh tags are provided in Chung et al. 2016. SMASh tags comprise a site-specific drug-inhibitable protease (e.g. a hepatitis C virus (HCV) nonstructural protein 3 (NS3) protease domain) and a degron (e.g. a NS4A peptide). A SMASh tag can be fused to a target protein separated by a cleavage sequence for the protease included in the SMASh tag. This leads to separation of the tag in the absence of an inhibitor of the protease (e.g. simeprevir, danoprevir, asunaprevir, or ciluprevir), and degradation of the target protein in the presence of the inhibitor of the protease. In embodiments using a SMASh tag, the SMASh tag may be separated from the rest of the reporter construct by a cleavage sequence forthe protease included in the SMASh tag, and the degradation ligand may be an inhibitor of the protease included in the SMASh tag. For example, the reporter construct may comprise a selection marker 10, a protease recognition sequence 12, and a degron 14 wherein the degron comprises a SMASh tag and is separated from the rest of the reporter construct by a further protease recognition sequence different from the protease recognition sequence 12, the further protease recognition sequence comprising a protease cleavage motif recognised by the protease included in the SMASh tag. The SMASh tag may comprise a hepatitis C virus (HCV) nonstructural protein 3 (NS3) protease domain. In such embodiments, a degradation ligand may be an inhibitor of NS3, such as e.g. simeprevir, danoprevir, asunaprevir, or ciluprevir. In embodiments, the degron is an NS4A peptide (e.g. the degron part of a SMASh tag, without the protease part). In such embodiments, a degradation ligand may not be used. In embodiments, the degron is a proteolysis targeting chimera (PROTAC) target protein or peptide. A PROTAC is a heterobifunctional small molecule comprising a ligand that binds a target protein, and a ligand that binds an E3 ubiquitin ligase, separated by a linker. The E3 ubiquitin ligase can be e g. pVHL, CRBN, Mdm2, beta-TrCP1 , DCAF15, DCAF16, RNF114 or c-IAP1. PROTACs comprising ligands binding do each of these E3 ubiquitin ligases have been previously described (see e.g. Li & Song, 2020). The target protein or peptide can be any protein or part thereof for which a ligand can be designed, including in particular any protein or part thereof for which PROTACs are known, such as e.g. BRD4, BTK, BCR-ABL, MCL1 , FLT-3, STAT3, BAF and AR. Examples of ligands targeting each of these are described in Li & Song, 2020. In embodiments in which the degron is a PROTAC target protein or peptide, the degradation ligand is a PROTAC. In embodiments, the degron is a HaloTag (e.g. HaloTag7, HaloTag2) protein and the degradation ligand is a HaloPROTAC. HaloTag is a modified bacterial dehalogenase that covalently reacts with hexyl chloride tags. A HaloPROTAC is a molecule comprising a E3 ubiquitin ligase ligand (e.g. a VHL or clAP ligand) linked to a chloroalkane moiety that can form a covalent bond with the HaloTag. HaloPROTACs that can be used for the degradation of any HaloTag? fusion protein (e.g. a reporter construct as described herein comprising a HaloTag? as degron) are available from Creative Biolabs (see e.g. www.creative-biolabs.com / protac / haloprotac.htm).
[0083] A protease recognition sequence (also referred to herein as a “protease cleavage sequence” or “protease cleavage site”) is an amino acid sequence that is specifically recognised by a protease, leading to cleavage of the protein or peptide comprising the recognition sequence. Any protease cleavage sequence known in the art may be used. In otherwords, the protease cleavage sequence may be a sequence that is specifically recognised by any protease known in the art. In the context of the present disclosure, the protease is an intracellular protease. An intracellular protease is a protease that can be present and active in an intracellular environment. The intracellular protease may be expressed by the cell in which the cleavage occurs or may be expressed by another cell and internalised by the cell in which the cleavage occurs. In the context of the present disclosure, a protease is typically an endoprotease. A protease recognition sequence may be selected from: a granzyme B (gzb) cleavage sequence, a gamma secretase cleavage sequence, or a caspase cleavage sequence. Thus, the protease may be selected from: granzyme B (gzb), gamma secretase and a caspase (such as e.g. caspase 3, caspase 7, caspase 8, and / or caspase 10). Reporters of caspase 8 / 10 activity are useful e.g. as indicators of Fas-mediated cell death. Reporters of gzb activity are useful e.g. as indicators of T cell reactivity. Reporters of gamma secretase activity are useful as indicators of pro-amyloid beta activity. Reporters of caspase 3 / 7 activity are useful as general indicators of apoptosis. The protease may be any protease described in Backes et al. 2005 and the protease cleavage sequence may be a sequence comprising the amino acid sequence of any cleavage site for the respective protease as described in Backes et al. 2005. A caspase may be selected from caspase 3, caspase 7, caspase 8, or caspase 10. In embodiments, the caspase is caspase 3 or 7 (caspase-3 / 7). In embodiments, the protease is gzb and the protease cleavage sequence is a sequence comprising amino acid sequence IEXD (SEQ ID NO: 23) where X can be any amino acid, VEXD (SEQ ID NO: 24) where X can be any amino acid, IGXD (SEQ ID NO: 25) where X can be any amino acid, or VGXD (SEQ ID NO: 26) where X can be any amino acid. In some such embodiments, X is selected from P, S, T, A, N or Q. For example, X may be P. In specific embodiments, the protease is gzb and the protease cleavage sequence is a sequence comprising amino acid sequence IEAD (SEQ ID NO: 1), IEPD (SEQ ID NO:17), orVGPD (SEQ ID NO: 2). In embodiments, the protease is caspase-3 / 7 and the protease cleavage sequence is a sequence comprising amino acid sequence DXXD (SEQ ID NO: 18) or DEXD (SEQ ID NO: 27) where X can be any amino acid. In some embodiments in which the protease cleavage sequence is a sequence comprising amino acid sequence DXXD, the first X is S, T, A, E, V, or Q, and the second X is V, P, I or T. In some embodiments in which the protease cleavage sequence is a sequence comprising amino acid sequence DEXD, X is V, P, I or T. In specific embodiments, the protease is caspase-3 / 7 and the protease cleavage sequence is a sequence comprising amino acid sequence DEVD (SEQ ID NO: 3). In embodiments, the protease is caspase-8 / 10 and the protease cleavage sequence is a sequence comprising amino acid sequence XEXD (SEQ ID NO: 28), where X can be any amino acid. In some such embodiments, the first X is P, E, D, A, I, L or V, and the second X is T, A, H, I, W or V. In embodiments, the protease is caspase-8 / 10 and the protease cleavage sequence is a sequence comprising amino acid sequence LQTD (SEQ ID NO: 4), LETD (SEQ ID NO: 29), VETD (SEQ ID NO: 30), LEHD (SEQ ID NO: 31), VEHD (SEQ ID NO: 32). In embodiments, the protease is gamma secretase. Gamma secretase is an intramembrane aspartyl protease complex. It is known to cleave the C99 fragment of amyloid precursor protein (APP), leading to the generation of extracellular Ap peptides. It also cleaves other substrates such as Notchl . In embodiments, the protease cleavage sequence is a sequence comprising amino acid sequence EDVGSNKGAIIGLMVGGVVIATVIVITLVMLKKK
[0084] (SEQ ID NO: 33) corresponding to E22-K55 of the C99 fragment of human APP, or a homologue or variant thereof (including e.g. a fragment thereof) that is recognised by gamma secretase. These amino acids have been shown to be sufficient for gamma secretase cleavage, see e.g. Yan et al. 2017. A variant of SEQ ID NO: 33 can be a sequence with at least 80%, at least 85%, at least 90%, at least 95%, at least 98% sequence identity to the sequence of SEQ ID NO:33, provided that the sequence is still recognised by gamma secretase. Human APP is available under Uniprot ID P05067. Mouse APP is available under Uniprot ID P12023. In embodiments, the protease cleavage sequence is a sequence comprising the transmembrane region of human Notch 1 , Notch2, Notch3 or Notch4, or a homologue thereof. Human Notchl is available as Uniprot ID P46531 , with a transmembrane region comprising amino acids 1736-1756, i.e.
[0085] FMYVAAAAFVLLFFVGCGVLL (SEQ ID NO: 34). Human Notch2 is available as Uniprot ID
[0086] Q04721 , with a transmembrane region comprising amino acids 1678-1698, i.e.
[0087] LLYLLAVAWIILFIILLGVI (SEQ ID NO: 35). Human Notch3 is available as Uniprot ID Q9UM47, with a transmembrane region comprising amino acids 1644-1664, i.e. LPLLVAGAVLLLVILVLGVMV (SEQ ID NO: 36). Human Notch4 is available as Uniprot ID Q99466, with a transmembrane region comprising amino acids 1448-1468, i.e. PVLCSPVAGVILLALGALLVL (SEQ ID NO: 37). In embodiments, the protease cleavage sequence is a sequence comprising amino acid sequence of SEQ ID NO: 34, 35, 36 or 37, or a homologue or variant thereof (including e.g. a fragment thereof) that is recognised by gamma secretase. A variant of SEQ ID NO: 34, 35, 36 or 37 can be a sequence with at least 80%, at least 85%, at least 90%, at least 95%, at least 98% sequence identity to the sequence of SEQ ID NO:34, 35, 36 or 37, provided that the sequence is still recognised by gamma secretase.
[0088] A protease recognition motif refers to a specific combination of amino acids that is recognised by a protease. In embodiments, the reporter comprises a protease recognition motif comprised in a peptide in which the protease recognition motif is expected to be exposed. Such a peptide may also be referred to herein as a protease recognition sequence. In other words, the term protease recognition sequence refers both to a specific amino acid sequence motif that is recognised by a protease, and to a portion of the reporter fusion proteins described herein that comprises said amino acid sequence motif. In embodiments, the reporter comprises a protease recognition motif comprised in a region forming an exposed loop in the reporter fusion protein. In embodiments, the exposed loop comprising the protease recognition sequence is a fragment of a BID protein, or a sequence comprising one or more mutations compared to a fragment of a BID protein (e.g. a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98% sequence identity to a fragment of a BID protein), provided that that sequence comprises a protease recognition motif. The protease recognition motif may be a granzyme B recognition motif. A BID protein refers to a mammalian BID protein or homologue thereof, such as e.g. Uniprot ID P55957 (human BID), or Uniprot ID P70444 (mouse BID). An example of an exposed loop comprising a granzyme B recognition sequence is shown on Fig. 1C. Fig. 1 C shows the predicted structure of the human protein BID (Uniprot ID P55957), with a protease recognition sequence highlighted between Leu47 and Ser78 (SEQ ID NO: 14), comprising a protease recognition motif starting at position 71 (although subsets thereof that comprise the granzyme B recognition motif can also be used). The mouse homologue sequence (Uniprot ID P70444) comprises a granzyme B recognition sequence starting at position 72, in an exposed loop comprising amino acids 45 to 78 (SEQ ID NO: 17). Thus, the reporter construct may comprise a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98% sequence identity to the sequence of amino acids 47 to 78 of human BID (SEQ ID NO: 14) or a sequence comprising at most 6, at most 5, at most 4, at most 3 or at most 2 mutations compared to the amino acid sequence of SEQ ID NO: 14, comprising a Granzyme B protease recognition motif (e.g. IEAD at positions 26 to 29 of SEQ ID NO: 14). In embodiments, the reporter comprises a fragment of the sequence provided as SEQ ID NO: 14 comprising at least the granzyme B recognition motif IEAD or an alternative granzyme B recognition motif such as IEPD. In embodiments, the reporter comprises a fragment of the sequence provided as SEQ ID NO: 14 comprising: (i) the granzyme B recognition motif IEAD or an alternative granzyme B recognition motif such as IEPD; and (ii) one or more flanking residues either side of the granzyme B recognition motif IEAD, such as e.g. 1 , 2 or 3 flanking residues C- terminal of the granzyme B recognition motif and 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10 or more residues N- terminal of the granzyme B recognition motif. Alternatively, the reporter construct may comprise a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98% sequence identity to the sequence of amino acids 45 to 78 of mouse BID (SEQ ID NO: 17) or a sequence comprising at most 6, at most 5, at most 4, at most 3 or at most 2 mutations compared to the amino acid sequence of SEQ ID NO: 17, comprising a Granzyme B protease recognition motif (e.g. IEPD at positions 28 to 31 of SEQ ID NO: 17). In embodiments, the reporter comprises a fragment of the sequence provided as SEQ D NO: 17 comprising at least the granzyme B recognition motif IEPD or an alternative granzyme B recognition motif such as IEAD. In embodiments, the reporter comprises a fragment of the sequence provided as SEQ ID NO: 17 comprising: (i) the granzyme B recognition sequence IEPD or an alternative granzyme B recognition motif such as IEAD; and (ii) one or more flanking residues either side of the granzyme B recognition motif IEPD, such as e.g. 1 , 2 or 3 flanking residues C-terminal of the granzyme B recognition motif and 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10 or more residues N-terminal of the granzyme B recognition motif. In embodiments, the reporter comprises the sequence provided as SEQ ID NO: 8, or a sequence comprising 1 , 2, 3, or 4 mutations compared to said sequence, the sequence comprising a granzyme B recognition motif such as IEAD (as per SEQ ID NO:8) or an alternative granzyme B recognition motif such as IEPD.
[0089] In embodiments, the exposed loop comprising the protease recognition sequence is a fragment of a PARP1 protein, or a sequence comprising one or more mutations compared to a fragment of a PARP1 protein (e.g. a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98% sequence identity to a fragment of a PARP1 protein), provided that that sequence comprises a protease recognition motif. The protease recognition motif may be a caspase 3 / 7 recognition motif. A PARP1 protein refers to a mammalian PARP protein or homologue thereof, such as e.g. Uniprot ID P09874 (human PARP1), or Uniprot ID P11103 (mouse PARP1). An example of an exposed loop comprising a caspase 3-7 recognition sequence is shown on Fig. 1E. Fig. 1 E shows the predicted structure of the human protein PARP1 (Unitprot ID P09874), with a protease recognition motif highlighted between V202 and D217 (provided as SEQ ID NO: 19), starting at position 21 1 (although subsets thereof that comprise the caspase recognition sequence can also be used). The mouse homologue sequence (Uniprot ID P11103) comprises a caspase 3 / 7 recognition sequence starting at position 21 1 , in an exposed loop comprising amino acids 202 to 217 (SEQ ID NO: 20). Thus, the reporter construct may comprise a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98% sequence identity to the sequence of amino acids 202 to 217 of human PARP1 (SEQ ID NO: 19) or a sequence comprising at most at most 3 or at most 2 mutations compared to the amino acid sequence of SEQ ID NO: 19, comprising a caspase 3 / 7 protease recognition motif (e.g. DEVD at positions 10 to 13 of SEQ ID NO: 19). In embodiments, the reporter comprises a fragment of the sequence provided as SEQ ID NO: 19 comprising at least the caspase 3 / 7 recognition motif DEVD or an alternative caspase 3 / 7 recognition motif such as DXXD. In embodiments, the reporter comprises a fragment of the sequence provided as SEQ ID NO: 19 comprising: (i) the caspase 3 / 7 recognition motif DEVD or an alternative caspase 3 / 7 motif such as DXXD; and (ii) one or more flanking residues either side of the caspase 3 / 7 recognition motif DEVD, such as e.g. 1 , 2 or 3 flanking residues C-terminal of the caspase 3 / 7 recognition motif and 1 , 2, 3, 4, 5, 6, 7, 8, or 9 residues N-terminal of the caspase 3 / 7 recognition motif. Alternatively, the reporter construct may comprise a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98% sequence identity to the sequence of amino acids 202 to 217 of mouse PARP1 (SEQ ID NO: 20) or a sequence comprising at most 3 or at most 2 mutations compared to the amino acid sequence of SEQ ID NO: 20, comprising a caspase 3 / 7 protease recognition motif (e.g. DEVD at positions 10 to 13 of SEQ ID NO: 20). In embodiments, the reporter comprises a fragment of the sequence provided as SEQ D NO: 20 comprising at least the caspase 3 / 7 recognition motif DEVD or an alternative caspase 3 / 7 recognition motif such as DXXD. In embodiments, the reporter comprises a fragment of the sequence provided as SEQ ID NO: 20 comprising: (i) the caspase 3 / 7 recognition sequence DEVD or an alternative caspase 3 / 7 recognition motif such as DXXD; and (ii) one or more flanking residues either side of the caspase 3 / 7 recognition motif DEVD, such as e.g. 1 , 2 or 3 flanking residues C-terminal of the caspase 3 / 7 recognition motif and 1 , 2, 3, 4, 5, 6, 7, 8, or 9 residues N- terminal of the caspase 3 / 7 recognition motif.
[0090] In a specific example, the coding sequence for any selection marker (e.g. a fluorescent marker, such as e.g. GFP, for example using a sequence such as that of SEQ ID NO: 7, mCherry - for example using a sequence such as that of SEQ ID NO: 50, or StayGold - for example using a sequence such as that of SEQ ID NO: 51) can be cloned in-frame with the coding sequence for a degron (e.g. mutant FKB12F36V, such as SEQ ID NO: 9) separated by the coding sequence for a specific protease cleavage sequence (e.g. a sequence comprising a protease recognition motif recognised specifically by Granzyme B, e.g. SEQ ID NO: 1 , 2, or 17, or a sequence comprising a motif recognised by caspase 3 / 7, e.g. SEQ ID NO: 3 or 18). In another example, a reporter may include a cell surface marker (e.g. RQR8 (e.g. as shown in SEQ ID NO: 52), CD34, CD19, NGFR (e.g. as shown in SEQ ID NO: 53), PD-L1) where the cell surface marker is expressed on the cell surface after protease cleavage enabling sorting by FACS or magnetic cell sorting (MACS, e.g. using cell separation beads). The cell surface marker may also function as an immune suppression protein. For example, a reporter may include a cell surface marker (e.g. PD-L1 ) where the cell surface marker is expressed on the cell surface after protease cleavage enabling sorting by FACS or magnetic cell sorting, and where the cell surface marker confers a survival advantage to the cells after protease cleavage in the presence of immune pressure. As another example, a reporter may include a cell survival protein such as BCL2, which is expressed on protease cleavage.
[0091] Suitably, the cell surface marker may be an RQR8 protein - for example as shown in SEQ ID NO: 52 or a variant with at least 80%, at least 85%, at least 90%, at least 95% of at least 99% thereto. Suitably, the cell surface marker may be an NGFR protein - for example as shown in SEQ ID NO: 53 or a variant with at least 80%, at least 85%, at least 90%, at least 95% of at least 99% thereto.
[0092] In a specific example, the coding sequence for any selection marker can be cloned in frame with the coding sequence for a protease cleavage sequence such as SEQ ID NOs: 8, 14, 17, 19 or 20 or fragments or homologues thereof as described above, and the coding sequence for a degron such as SEQ ID NO: 9. For example, the reporter construct can comprise a sequence coding for a selection marker, upstream of the sequence of SEQ ID NO: 21 (combining a protease cleavage sequence comprising a gzmb recognition motif, and a FKB12F36V peptide). As another example, the reporter construct can comprise a sequence coding for a selection marker, upstream of the sequence of SEQ ID NO: 22 (combining a protease cleavage sequence comprising a casp3 / 7 recognition motif, and a FKB12F36V peptide).
[0093] The reporter construct and reporter fusion protein are designed so that the degron is accessible to the degradation ligand, thereby enabling inducible degradation of the degron containing fusion protein as described herein. Design considerations to enable binding of the degradation ligand to the degron include the number of amino acids between the selection marker and the degron within the reporter fusion protein. Suitably, at least 10, at least 12, at least 14, at least 16, at least 18, at least 20, at least 30, or at least 40 amino acids may be included between the selection marker and the degron in the reporter fusion protein. Inducible degradation of the reporter fusion protein may be determined by any suitable means known in the art, for example, as described herein (see Example 1).
[0094] The reporter construct can be included in an expression vector. Any expression vector suitable for expression of a fusion protein in the target cells can be used. In embodiments, the vector is a viral vector, such as a lentiviral vector. In embodiments, the vector comprises the reporter construct (i.e. a nucleic acid sequence encoding a reporter fusion protein as described herein) and one or more additional constructs. For example, the one or more additional constructs can comprise one or more of: an antigen expression construct, a selection transduction selection marker (e.g. an antibiotic resistance cassette), one or more promoter sequences (e.g. a CMV promoter, SFFV promoter or hPGK promoter) and one or more enhancer sequences. An example of a vector suitable for use in the present disclosure is shown on Fig. 11. This comprises a sequence encoding the reporter fusion protein as described herein (including a selection marker sequence 10, a protease recognition sequence 12 and a degron sequence 14), a promoter sequence 18, an enhancer sequence 19, and a plurality of antibiotic resistance genes 17 and associated promoters 18.
[0095] A “target cell’’ refers to a cell that expresses a reporter fusion protein as described herein. The target cell may be any cell in which a protease associated cellular event is to be identified. The protease associated cellular event may be targeting of the cell by a cytotoxic T cell. In such embodiments, the target cell may be an antigen presenting cell. In some such embodiments, the protease recognition sequence is a granzyme B recognition sequence. The protease associated cellular event may be apoptosis. In such embodiments, the protease recognition sequence may be a caspase recognition sequence. In some such embodiments, the protease recognition sequence is a caspase-3 / 7 or caspase-8 / 10 recognition sequence. The target cell may have been modified to knockout one or more pro-apoptotic proteins or comprise a compound that inhibits one or more pro-apoptotic proteins. The one or more pro-apoptotic proteins can be selected from: BIM, BAX, BIK, BAD, BAK, BID, NOXA, BMF, and HRK. This enables survival of the target cell even in the presence of proapoptotic I cytotoxic signals, which increases the likelihood of recovering positive target cells.
[0096] A “triplet” or “MHC-antigen-TCR triplet”, “MHC-peptide-TCR triplet” or “MHC-antigen-T cell” triplet refers to a set of molecules or the sequences thereof comprising: a peptide (also referred to as “antigen peptide”), an MHC molecule capable of presenting the peptide, and a T cell receptor (TCR) capable of recognising (i.e. binding to) the peptide in the context of the MHC molecule (pMHC). An immunogenic triplet (or “activated triplet”) refers to a triplet that is such that the TCR binds to the antigen peptide in the context of the MHC molecule, where the binding results in activation of the T cell. This recognition process underlines the triggering of an immune reaction, also referred to as “cellular immune reaction”. The terms “antigen”, “peptide” and “antigen peptide” are used interchangeably to referto a peptide that is potentially immunogenic. Thus, such a peptide can also be referred to as a candidate antigen peptide. Figure 2A illustrates the process of antigen recognition and how this can be detected in embodiments of the disclosure. An MHC molecule 3 is expressed on the surface of a cell 2 (which is referred to herein as an “antigen presenting cell” (APC) or “target cell”, which can be a professional antigen presenting cell or any cell presenting intracellular antigens, such as e.g. a cancer cell, cell infected by a virus, or a modified immortalised cell, such as a cell from a cancer cell line, here illustrated as a k562 cell). The MHC molecule displays a peptide 6. The complex formed by the MHC molecule 3 and the peptide 6 is recognised by a T cell receptor 5 expressed on the surface of a T cell 4. Upon engagement between the T cell receptor s and a cognate peptide-MHC complex, additional engagement of coreceptors 5a occurs, the T cell is activated and secretes perforin (PFN) and granzyme B (GzmB). Perforin creates pores in the antigen presenting cell 2, leading to the diffusion of GzmB inside of these cells. The antigen presenting cell 2 expresses a reporter fusion protein as described herein, which in this case is illustrated as a reporter as shown on Figure 1 B. In the absence of granzyme B, i.e. in the absence of engagement of the antigen presenting cell with an activated T cell, the reporter fusion protein is intact and detectable in the absence of the dTag-13 molecule (i.e. the antigen presenting cells are positive for the selection marker forming part of the reporter fusion protein), and degraded in the presence of the dTag-13 molecule (i.e. the antigen presenting cells are negative for the selection marker forming part of the reporter fusion protein). In the presence of engagement of the antigen presenting cell with an activated T cell, the reporter fusion protein is cleaved by GzmB secreted by the activated T cell, and the selectable maker (here illustrated as GFP) is cleaved from the degron. As a result, the antigen presenting cells are positive for the selection marker forming part of the reporter fusion even in the presence of the dTag-13 molecule. Note that the system can also function without a conditional degron, in which case the antigen presenting cells are simply negative for the selection marker in the absence of activated T cell engagement, and positive for the selection marker in the presence of activated T cell engagement. However, the use of a conditional degron advantageously enables confirmation of successful transformation of the cells to express the reporter fusion protein, thereby reducing the risk of false negatives. Figure 2B illustrates how process of antigen recognition can be probed in a high throughput, flexible manner using assays of the disclosure. A modified immortalised cell 2 is provided which does not express any genomic copies of MHC class I and / or class II molecules. An expression vector 8 is introduced in the cell 2, which comprises the coding sequence for a single MHC allele, enabling the expression of an MHC molecule 3 on the cell surface. The MHC molecule can be a class I MHC molecule or a class II MHC molecule. The MHC allele may have been introduced into the cells using any method known in the art. In embodiments, the MHC molecule is introduced in the cell 2 using an expression vector 8 which includes a construct comprising the coding sequence for a single MHC allele, and a pair of priming sites on either sides of the coding sequence. The expression vector may further encode a polyA sequence. The priming sites enable identification of the MHC allele by single cell RNA sequencing of recovered antigen presenting cells. The priming sites may be specific to the MHC allele, i.e. they may be distinct from any other priming site associated with another MHC allele used in the same experiment. This enables confident identification of the MHC allele expressed by cell 2 in any recovered cell in experiments where multiple MHC alleles are tested simultaneously using a plurality of populations of cells 2 that each express a respective MHC allele. The polyA sequence enables polyT based capture of the RNA expressed from the construct and comprising the MHC allele coding sequence, for identification of the MHC allele by single cell RNA sequencing. In the illustrated embodiment, the expression vector 8 (or a separate expression vector) also comprises a construct 9 comprising the coding sequence for an antigen, as illustrated herein a neoantigen transcript. The construct 9 may also advantageously contain priming sites and a sequence encoding a poly-A tail to enable identification of the antigen by single cell RNA sequencing of recovered doublets. As explained above, a polyA tail can be used for polyT based capture of the RNA expressed from the construct and comprising the antigen, for identification of the antigen by single cell RNA sequencing. The priming sites enable specific amplification of the sequence of the antigen (after polyT based capture if this step is performed). The priming sites may be universal priming sites, i.e. sites that are complementary to universal primers. A universal primer and corresponding priming site may refer to sites / primers that have sequences that are present in a plurality or all of a set of constructs 9 each including a coding sequence for a different antigen. Thus, the same primers can be used to amplify and recover any antigen present in a recovered doublet. The priming sites (and corresponding primers) may be designed to have a sequence that is not expected to occur in the immortalised cell 2 in the absence of the construct 9. For example, the priming sites may be designed to have a sequence that does not occur in human cells. In the illustrated embodiment, the expression vector 8 (or a separate expression vector) also comprises a construct 7 comprising the coding sequence for a granzyme B activated reporter fusion protein as described herein. Any expression vector that enables antigen processing may be used to introduce construct 9 and / or construct 7 into the cells. For example, the expression vector may be a viral expression vector, such as a lentiviral expression vector. Lentiviral expression vectors advantageously enable expression of very large antigen libraries. In other embodiments, the expression vector may be a minigene or tandem minigenes. Tandem minigenes have been shown to achieve high expression in various neoantigen expression systems, but are not as easily amenable to large libraries. In other embodiments, the neoantigen may be provided as a peptide that is incubated with the immortalised cell expressing the chosen MHC molecule. This is less advantageous as it does not allow identification of the neoantigen that is involved in an antigen recognition event unless single antigens are included in each coculture. Further, endogenous processing of an antigen is believed to be an important element of reliable recognition of an antigen, such that endogenous expression of the antigen is likely to enable more reliable detection of whether a peptide sequence truly can be processed, presented and recognised by a T cell. The cell 2 is cocultured with a T cell 4. This is typically a T cell that is part of a complex population of T cells, for example a primary T cell population or a T cell population derived from such a T cell population using one or more specific and / or non-specific expansion steps. The T cell natively expresses a specific T cell receptor 5. In the illustrated embodiment, the T cell is a CD8+ T cell. However, the T cell may be a CD4+ T cell or a CD8+ T cell. Thus, the T cell can be part of a population of T cells comprising CD4+ T cells and / or CD8+ T cells. The CD4+ T cells may be cytotoxic T cells. Co-receptors (including CD8 which binds to HLA) and costimulation molecules 5a are also expressed by the T cell (although they are not specific to an antigen). If the peptide 6 expressed by the cell 2 is presented in the context of the MHC molecule 3 (resulting in a pMHC complex), and if further this pMHC complex is recognised by the T cell 4, a doublet of physically interacting cells (cell 2 and cell 4) forms. This may further lead to activation of the T cell, in which case the peptide 6 is deemed immunogenic in the context of the MHC molecule 3. As explained above, activation of the T cell leads to secretion of perforin and granzyme B, and eventually apoptosis of the target cell (thus it is also possible to use a caspase-activated reporter fusion protein instead of a granzyme B activated fusion protein). In the context of identifying antigen-T cell recognition events, the use of a granzyme B activated reporter is advantageous as granzyme B secretion is an early indicator of T cell activation, and therefore can be detected quickly, prior to significant cell death has occurs (at which point it is no longer possible to recover the antigen presenting cells and associated immunogenic antigens). The population of T cells can comprise naive CD4+ T cells that have been pre-expanded in the presence of CD3 / CD28 activators and IL2, to obtain cytotoxic CD4+ T cells. Cytotoxic CD4+ T cells may produce similar markers of T cell activation as CD8+ T cells. When the T cell has been activated, the doublet of cells becomes a positive doublet, i.e. a doublet of cells comprising an antigen presenting cell that is positive for the selection marker forming part of the (now cleaved) reporter fusion protein, as illustrated on Figure 2C. Positive doublets can be recovered by any cell sorting known in the art that is suitable for sorting using the particular selection marker used. In the embodiment illustrated on Figure 2C, the selection marker is GFP (although any fluorescent protein can be used in its place), such that the positive doublets can be recovered by fluorescence activated cell sorting (FACS). These can then be subject to single cell sequencing (or even bulk sequencing if it is not a requirement to identify individual antigen-TCR pairs), to identify the antigen and / or T cell receptor that were involved in the APC-T cell engagement.
[0097] A “doublet” refers to a pair of physically interacting cells comprising a cell presenting an antigen in the context of an MHC molecule (i.e. a cell comprising a pMHC complex on its cell surface), and a T cell expressing a TCR that recognises the pMHC complex on the surface of the presenting cell. Doublets can be activated (also referred to as true positive doublets) or not activated (also referred to as false positive doublets). Note that false positive doublets do not necessarily only contain doublets that result from unspecific interactions or interactions that cannot lead to activation. Indeed, some doublets may not be activated at the time of detection because they have not had enough time to upregulate expression of markers of activation. Thus, false positive doublets may also be referred to as “unconfirmed positives”. An activated doublet comprises an activated T cell. A doublet that is not activated comprises a T cell that is not activated. Activation of the T cell refers to the triggering of intracellular signalling events following pMHC-TCR binding. T cell activation can lead to one or more of: the secretion of cytokines, cell proliferation, secretion of effector molecules such as granzyme B and perforins, etc. T cell activation is associated with changes in RNA expression that are detectable through e.g. single cell RNA sequencing. The present inventors have demonstrated (data not shown) that a significant proportion of doublets in doublets sorted after T cell- antigen presenting cell co-culture are not in fact activated, and that it is therefore beneficial to further filter doublets to select (truly) activated doublets. The reporters of the present disclosure enable such selection.
[0098] An antigen presenting cell refers to a cell that displays an antigen on its cell surface, in the context of a MHC molecule. Embodiments of the present disclosure makes use of immortalised cells as antigen presenting cells. These cells may be obtained from any immortalised cell line, preferably any cancer cell line known in the art. In embodiments, the cells are cancer cells. The cells can be obtained from a cell line that is HLA-null, for example through B2M mutation. This enables the cells to be modified to be HLA monoallelic. Advantageously, the cells may be from a leukaemia cell line, such as k562 or 721 .221. 721.221 is a human HLA-negative B-lymphoblastoid cell line. K562 is a chronic myelogenous leukemia cell line that is negative for class-l HLA expression (HLA-null) but positive for class-l antigen processing and presentation pathway. Therefore, the cell line can be used to express any antigen, for example as a minigene, in the context of a chosen HLA allele inserted into the cells. K562 is a leukemia cell line that is very well characterised genetically and has been previously used in immunopeptidomics settings. The cell line may have been genetically modified in a number of ways. The cells may have been genetically modified to express or overexpress an anti-apoptotic factor. An anti-apoptotic factor is typically a protein that, when expressed by the cell, counters or reduces the effect of pro-apoptotic signalling in the cell. For example, the anti-apoptotic factor is selected from: BCL2, BCL-XL, MCL-1 , BFL-1 , BCL-W and BCL2L10.
[0099] MHC (major histocompatibility complex) molecules are cell-surface proteins encoded by the human leukocyte antigen (HLA) gene complex, and which are an important part of the adaptive immune system. MHC molecules are typically classified as “class I” or “class II”. Class I MHC molecules present peptides from inside the cells for recognition by T cell receptors as will be explained further below. Class I MHC molecules are normally expressed on the surface of all cells. The peptides presented are typically produced from digested proteins produced in the proteasomes, and are typically about 8-1 1 amino acids in length. There are 3 types of MHC class I molecules (A, B and C), each encoded by a separate gene. Class II MHC molecules present antigens from outside the cells for recognition by T cells. Class II MHC molecules are primarily found on antigen-presenting cells such as dendritic cells, mononuclear phagocytes, some endothelial cells, thymic epithelial cells and B cells. There are 6 types of MHC class II molecules: DP, DM, DOA, DOB, DQ and DR, each encoded by a separate gene. The HLA locus is highly polymorphic and therefore many different alleles exist for each gene. The process of HLA typing refers to determining which alleles of each of one or more HLA genes is present in a sample or subject. Methods for HLA typing are known in the art and include e.g. flow cytometry-based methods and methods based on sequencing data such as Polysolver (Shukla et al. 2015) and OptiType (Szolek et al. 2014). The term “MHC sequence” as used herein refers to the amino acid sequence of an MHC molecule or part thereof, or a nucleic acid sequence coding for such an amino acid sequence. In the context of the present disclosure, the MHC molecule may be a class I MHC molecule or a class II MHC molecule.
[0100] The term “TCR sequence” as used herein refers to the amino acid sequence of a T cell receptor or part thereof, or a nucleic acid sequence coding for such a sequence. A T cell receptor is a membrane anchored protein expressed on the surface of T cells. A T cell receptor comprises a pair of protein chains that together form binding moiety that recognises a cognate antigen. The TCR chains are expressed in a complex with constant T cell coreceptor chains CD3 (illustrated as reference numeral 7 in Figure 1 ), comprising a CD3y chain, a CD36 chain, and two CD3E chains in mammals. The constant chains associate with the T cell receptor and the constant -chain to form the TCR complex, which together is able to generate a signal upon antigen binding to the T cell receptor. As illustrated on Figure 2B, the TCR 5 is a heterodimeric protein, comprising two highly variable chains, the a and chains (in the majority of T cells), or the alternative y and 5 chains (in a minority of T cells). Each chain comprises two extracellular domains: a variable region (or variable domain) and a constant region (or constant domain, proximal to the cell membrane), a transmembrane region and a short cytoplasmic tail. The variable regions together bind to a peptide (antigen) 6, within the context of a MHC (major histocompatibility complex) molecule 3 in the case of ap TCRs. Each variable domain contains three hypervariable regions referred to as the complementarity-determining regions (CDRs, respectively referred to as CDR1 , CDR2 and CDR3 on each of the chains), which together form an antigen binding site. A TOR sequence may comprise the complete sequence of one or both chains of a TCR, or a part of one or both chains.
[0101] A “sample” as used herein may be a cell or tissue sample, a biological fluid, an extract (e.g. a DNA extract obtained from the subject), from which genomic and / or transcriptomic material can be obtained for genomic and / or transcriptomic analysis, such as genomic sequencing (e.g. whole genome sequencing, whole exome sequencing) or RNA sequencing (also referred to as “RNAseq” or “RNA-seq”). The sample may be a cell, tissue or biological fluid sample obtained from a subject (e.g. a biopsy). Such samples may be referred to as “subject samples”. In particular, the sample may be a blood sample, or a tumour sample, or a sample derived therefrom, such as e.g. by cell purification (when the sample comprises cells, such as e.g. T cells) and / or DNA or RNA extraction (when the sample is used for genomic or transcriptomic analyses, such as e.g. for the purpose of identifying neoantigens to be screened). The sample may be a cell or tissue culture sample. As such, a sample as described herein may refer to any type of sample comprising cells or genomic and / or transcriptomic material derived therefrom, whether from a biological sample obtained from a subject, or from a sample obtained from e.g. a cell line. As used herein, a “subject” is preferably a mammalian subject (such as e.g. a cat, dog, horse, donkey, sheep, pig, goat, cow, mouse, rat, rabbit or guinea pig), preferably a human. Further, the sample may be transported and / or stored, and collection may take place at a location remote from the immune assay location (where the steps of co-culture and doublet isolation may be performed at the same or a separate location from any subsequent steps such as sequence data acquisition). Further, any computer-implemented method steps described herein (such as e.g. sequence data analysis) may take place at a location remote from the sample collection location and / or remote from the sequence data acquisition (e.g. sequencing) location (e.g. the computer-implemented method steps may be performed by means of a networked computer, such as by means of a “cloud” provider).
[0102] A “tumour sample” refers to a sample derived from or obtained from a tumour. Such samples may comprise tumour cells and normal (non-tumour) cells. The normal cells may comprise immune cells (such as e.g. lymphocytes), and / or other normal (non-tumour) cells (e.g. stromal cells). The lymphocytes in such mixed samples may be referred to as “tumour-infiltrating lymphocytes” (TIL). A tumour may be a solid tumour or a non-solid or haematological tumour. A tumour sample may be a primary tumour sample, tumour-associated lymph node sample, or a sample from a metastatic site from the subject. A sample comprising tumour cells or genetic material derived from tumour cells may be a bodily fluid sample. Thus, the genetic material derived from tumour cells may be circulating tumour DNA or tumour DNA in exosomes. Instead or in addition to this, the sample may comprise circulating tumour cells.
[0103] A “normal sample”, “healthy sample” or “germline sample” refers to a sample that is assumed not to comprise tumour cells or genetic material derived from tumour cells. A germline sample may be a blood sample, a tissue sample, or a purified sample such as a sample of peripheral blood mononuclear cells from a subject. Similarly, the terms “normal”, “germline” or “wild type” when referring to sequences or genotypes refer to the sequence I genotype of cells other than tumour cells. A germline sample may comprise a small proportion of tumour cells or genetic material derived therefrom, and may nevertheless be assumed, for practical purposes, not to comprise said cells or genetic material. In other words, all cells or genetic material may be assumed to be normal and / or sequence data that is not compatible with the assumption may be ignored.
[0104] The terms “tumour-specific mutation”, “somatic mutation” or simply “mutation” are used interchangeably and refer to a difference in a nucleotide sequence (e.g. DNA or RNA) in a tumour cell compared to a healthy cell from the same subject. The difference in the nucleotide sequence can result in the expression of a protein which is not expressed by a healthy cell from the same subject. For example, a mutation may be a single nucleotide variant (SNV), multiple nucleotide variant (MNV), a deletion mutation, an insertion mutation (together, “indel mutation”), a translocation, a missense mutation, a translocation, a fusion, a splice site mutation, or any other change in the genetic material of a tumour cell. A mutation may result in the expression of a protein or peptide that is not present in a healthy cell from the same subject. Mutations may be identified by exome sequencing, RNA-sequencing, whole genome sequencing and / or targeted gene panel sequencing and or routine Sanger sequencing of single genes, followed by sequence alignment and comparing the DNA and / or RNA sequence from a tumour sample to DNA and / or RNA from a reference sample or reference sequence (e.g. the germline DNA and / or RNA sequence, or a reference sequence from a database). Suitable methods are known in the art. An “indel mutation" refers to an insertion and / or deletion of bases in a nucleotide sequence (e.g. DNA or RNA) of an organism. Typically, the indel mutation occurs in the DNA, preferably the genomic DNA, of an organism. In embodiments, the indel may be from 1 to 100 bases, for example 1 to 9, 1 to 50, 1 to 23 or 1 to 10 bases. An indel mutation may be a frameshift indel mutation. A frameshift indel mutation is a change in the reading frame of the nucleotide sequence caused by an insertion or deletion of one or more nucleotides. Such frameshift indel mutations may generate a novel openreading frame which is typically highly distinct from the polypeptide encoded by the non-mutated DNA / RNA in a corresponding healthy cell in the subject.
[0105] An antigen peptide refers to a peptide that is capable of binding to an MHO molecule and interact with a TCR receptor in the context of an MHC molecule to elicit an immune response. The term “peptide” as used herein encompasses an antigen peptide and a peptide that is a candidate antigen peptide, i.e. a peptide for which immunogenicity is to be tested for example as described herein, or a fragment thereof. Antigen peptides may be synthesised using methods which are known in the art. The term "peptide" is used in the normal sense to mean a series of residues, typically L-amino acids, connected one to the other typically by peptide bonds between the a-amino and carboxyl groups of adjacent amino acids. The term includes modified peptides and synthetic peptide analogues. Antigen peptides may be neoantigens.
[0106] A “neoantigen” (or “neo-antigen”) is an antigen that arises as a consequence of a mutation within a cancer cell. Thus, a neoantigen is not expressed (or expressed at a significantly lower level) by normal (i.e. non-tumour) cells. A neoantigen may be processed to generate distinct peptides which can be recognised by T cells when presented in the context of MHO molecules. As described herein, neoantigens may be used as the basis for cancer immunotherapies. References herein to "neoantigen" are intended to include also peptides derived from neoantigens. The term "neoantigen” as used herein is intended to encompass any part of a neoantigen that is immunogenic.
[0107] An “antigenic” molecule as referred to herein is a molecule which itself, or a part thereof, is capable of stimulating an immune response, when presented to the immune system or immune cells in an appropriate manner. The binding of an antigen to a particular MHC molecule (encoded by a particular HLA allele) results in the antigen being presented by said MHC molecule on the cell surface, a, necessary but not sufficient condition for immunogenicity. Immunogenicity further requires recognition of the peptide-MHC complex by a T cell receptor. As used herein a “candidate antigen” refers to a peptide or sequence thereof that is potentially immunogenic, the immunogenicity of which has not yet been verified. The present disclosure provides methods to determine whether a candidate antigen is immunogenic, i.e. a bona fide antigen. The term antigen as used herein specifically encompasses antigens that arise as a consequence of a mutation within a cancer cell (i.e. neoantigen).
[0108] The binding of a neoantigen to a particular MHC molecule (encoded by a particular HLA allele) and / or the presentation of the neoantigen by the MHC molecule may also be predicted using methods which are known in the art. For example, MHC binding of neoantigens may be predicted using the netMHCpan4 (Jurtz et al. 2017) algorithm. A candidate neoantigen that has been predicted to bind to or be presented by a particular MHC molecule may be considered to be more likely to be presented by said MHC molecule on the cell surface. However, these methods are only predictive to a certain extent, and cannot determine whether a peptide-MHC complex, even if it was formed in vitro or in vivo, would in fact be recognised by any TCR, let alone a TCR present in a particular patient or subject. Indeed, immunogenicity is further believed to require interaction between the neoantigen with a T cell receptor (TCR) present in the subject, in the context of an MHC molecule present in the subject. The binding of peptides or peptide-MHC complexes to T cell receptors can be predicted using methods which are known in the art, such as e.g. PMTnet (Lu et al. 2021) and Imrex (Moris et al., 2020) and the method described in application PCT / EP2024 / 057046. Such methods may be used to select candidate neoantigens for testing using the methods described herein. However, all such predictive methods require knowledge of the TCR sequences to be screened, and are only predictive to a certain extent. Thus, assays that can positively verify immunogenicity in the context of a particular sample or patient are still needed. A neoantigen may be a candidate neoantigen that has been verified to be immunogenic using an assay as described herein. Further, the methods and products described herein can be used to generate data for training the machine learning models used in such methods.
[0109] A neoantigen peptide is a peptide that is encoded by a sequence comprising a cancer-specific mutation. The neoantigen peptide may comprise the cancer cell specific mutation (e.g the non- silent amino acid substitution encoded by a single nucleotide variant (SNV)) at any residue position within the peptide. By way of example, a peptide which is capable of binding to an MHC class I molecule is typically 7 to 13 amino acids in length. In embodiments, neoantigen peptides may be from 7 to 15, such as 8 to 13 amino acids in length. The peptides may be 7, 8, 9, 10, 1 1 , 12, 13, 14 or 15 amino acids long. In embodiments, 15 amino acids long peptides are designed as a set of overlapping sequences with 11 amino acid overlaps (i.e. a “jump” of 4 amino acids from the start of one sequence to the start of the next) each including at least one amino acid that results from the presence of a tumour specific mutation, to provide an overlapping peptide pool. In embodiments, longer peptides, for example 15-31-mers, may be used. In such embodiments, the mutation (i.e. any one or more amino acids resulting from the presence of a tumour specific mutation) may be at any position, for example at the centre of the peptide, e.g. at positions 7, 8, 9, 10, 11 , 12, 13, 14, 15 or 16. Such peptides can also be used to stimulate both CD4 and CD8 cells to recognise neoantigens. For examples, longer peptides, such as peptides that are 27, 28, 29, 30 or 31 amino acids long, may be used to stimulate both CD4+ and CD8+ cells. The mutation may be present at any residue position(s) within the peptide. Thus, the peptide may comprise one or more amino acids that are not present in a corresponding native peptide expressed by a healthy cell at any residue position(s) within the peptide.
[0110] A “clonal neoantigen” (also sometimes referred to as “truncal neoantigen”) is a neoantigen that results from a mutation that is present in essentially every tumour cell in one or more samples from a subject (or that can be assumed to be present in essentially every tumour cell from which the tumour genetic material in the sample(s) is derived). Similarly, a “clonal mutation” (sometimes referred to as “truncal mutation”) is a mutation that is present in essentially every tumour cell in one or more samples from a subject (or that can be assumed to be present in essentially every tumour cell from which the tumour genetic material in the sample(s) is derived). Thus, a clonal mutation may be a mutation that is present in every tumour cell in one or more samples from a subject. A “sub-clonal” neoantigen is a neoantigen that results from a mutation that is present in a subset or a proportion of cells in one or more tumour samples from a subject (or that can be assumed to be present in a subset of the tumour cells from which the tumour genetic material in the sample(s) is derived). Similarly, a “sub-clonal” mutation is a mutation that is present in a subset or a proportion of cells in one or more tumour samples from a subject (or that can be assumed to be present in a subset of the tumour cells from which the tumour genetic material in the sample(s) is derived). A neoantigen or mutation may be clonal in the context of one or more samples from a subject while not being truly clonal in the context of the entirety of the population of tumour cells that may be present in a subject (e.g. including all regions of a primary tumour and metastasis). Thus, a clonal mutation may be “truly clonal” in the sense that it is a mutation that is present in essentially every tumour cell (i.e. in all tumour cells) in the subject. This is because the one or more samples may not be representative of each and every subset of cells present in the subject. Thus, within the context of the present disclosure, a “clonal neoantigen” or “clonal mutation” may also be referred to as a “ubiquitous neoantigen” or “ubiquitous mutation”, to indicate that the neoantigen is present in essentially all tumour cells or all tumour samples that have been analysed, but may not be present in all tumour cells that may exist in the subject. The terms “clonal” and “ubiquitous” are used interchangeably unless context indicates that reference to “true clonality” was intended. The wording “essentially every tumour cell” in relation to one or more samples or a subject may refer to at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94% at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the tumour cells in the one or more samples or the subject. Nevertheless, a neoantigen / mutation that is identified as likely to be clonal (or “ubiquitous”) may be considered likely to be truly clonal, or at least more likely to be truly clonal than a neoantigen / mutation that is identified as unlikely to be clonal. Further, the confidence in the probability that a clonal neoantigen / mutation identified in a subject is truly clonal increases when the sample(s) used to identify the clonal neoantigen / mutation capture a more complete picture of the genetic diversity of the tumour (e.g. by including a plurality of samples from the subject, such as e.g. samples from different regions of the tumour, and / or by including samples that inherently capture a diversity of tumour cells such as e.g. ctDNA samples). Conversely, a neoantigen / mutation that is identified as unlikely to be clonal is unlikely to be truly clonal, because the identification that the neoantigen / mutation is unlikely to be clonal indicates that even in the restricted view afforded by the sampling process, there is evidence that the neoantigen / mutation is not present in all tumour cells. Thus, the process of identifying clonal neoantigens / mutations may be seen as prioritising which candidate neoantigens / mutations are most likely to be clonal, based on the restricted view of the clonal structure of the subject’s tumour available from the one or more samples. Methods to identify clonal neoantigens are known in the art and include the methods described in WO 2022 / 207925, WO 2016 / 16174085, and McGranahan et al. (2016). Clonal neoantigens may in particular be identified using a method as described in WO 2022 / 207925.
[0111] As used herein, the terms “computer system” includes the hardware, software and data storage devices for embodying a system or carrying out a method according to the above-described embodiments. For example, a computer system may comprise a central processing unit (CPU) and / or a graphic processing unit (GPU), input means, output means and data storage, which may be embodied as one or more connected computing devices. A computer system may comprise a display or comprises a computing device that has a display to provide a visual output display. The data storage may comprise RAM, disk drives or other non-transitory computer readable media. The computer system may include a plurality of computing devices connected by a network and able to communicate with each other over that network. It is explicitly envisaged that computer system may consist of or comprise a cloud computer.
[0112] As used herein, the term “computer readable media” includes, without limitation, any non-transitory medium or media which can be read and accessed directly by a computer or computer system. The media can include, but are not limited to, magnetic storage media such as floppy discs, hard disc storage media and magnetic tape; optical storage media such as optical discs or CD-ROMs; electrical storage media such as memory, including RAM, ROM and flash memory; and hybrids and combinations of the above such as magnetic / optical storage media.
[0113] Applications
[0114] The methods described herein find applications in any context in which it is desirable to detect cells in which a specific protease is active. One such situation is in the context of recognition of antigens by T cells. Examples of this will be described by reference to Figures 3A and 3B. Another such situation is in the context of detecting apoptosis or the effect of one or more compounds, compositions, cell populations or perturbations on apoptosis or gamma secretase activity. Examples of this will be described by reference to Figure 4.
[0115] Thus, the present disclosure provides methods for identifying a combination of an MHC allele, antigen and T cell (or T cell receptor, TOR) triplet, where the triplet is immunogenic, as well as methods of identifying any parts of an immunogenic triplet (e.g. immunogenic antigen, immunogenic MHC-antigen complex, TOR recognising an antigen, TCR-cognate antigen pair, MHC context in which an antigen is immunogenic). An illustrative method will be described by reference to Figure 3A. In the embodiments described by reference to Figure 3A (and Figure 3B), the protease recognition sequence used in the reporter fusion protein is typically a granzyme B recognition sequence. The selection marker may be any selection marker described herein, and the degron may be any degron described herein.
[0116] At step 212, a population of antigen presenting cells is obtained which is a population of cells modified to express a reporter fusion protein as described herein. This may comprise introducing an expression vector into the cells, the expression vector comprising a nucleic acid sequence encoding the receptor fusion protein. In embodiments, the population of antigen presenting cells may also be modified to express a selected HLA allele. This may comprise obtaining a sample comprising cells of a MHC monoallelic cell line, or obtaining a sample comprising cells of a MHC KO cell line and introducing the coding sequence of the selected MHC molecule into these cells. Alternatively, the cells obtained may be HLA-null cells that have been genetically modified to express the selected MHC allele. In any embodiment, the MHO allele may be a class I allele or a class II allele. In embodiments, the cells obtained at step 212 are immortalised cells from a MHC monoallelic cell line. In any embodiment, the antigen presenting cells may be immortalised cells, for example from a cancer cell line, such as a leukaemia cell line (e.g. a k562 cell line, which may be a native k562 cell line or a derived cell line that has been modified to be MHC monoallelic).
[0117] At step 214 the population of antigen presenting cells is modified to comprise cells presenting one or more of a plurality of candidate peptides on their cell surface in complex with the MHC molecule encoded by the selected HLA allele. This may comprise incubating the population of antigen presenting cells in the presence of the one or more candidate peptides. Alternatively, this may comprise introducing an expression vector in cells of the population of antigen presenting cells (for example by contacting the population of antigen presenting cells with a library of expression vectors in conditions suitable for introducing the expression vectors into the cells), the expression vector comprising a sequence encoding for one or more of the candidate peptides. This may be combined with step 212. In other words, the same expression vector may be used to introduce a sequence encoding for one or more candidate peptides / antigens and a sequence encoding a reporter fusion protein as described herein. An expression vector used at step 212 or step 214 may be a viral vector, such as a lentiviral expression vector, or a minigene (e g. a tandem minigene when coding sequences for multiple peptides are introduced in a single cell). The plurality of candidate peptides together forms a library of candidate peptides to be tested. Thus, a plurality of expression vectors each including the coding sequence for a candidate peptide (or multiple candidate peptides, although single candidate peptides advantageously make it simple to identify the peptide involved in the antigen recognition event underlying any single doublet) also form a library of vectors encoding the library of candidate peptides. The library of candidate peptides may comprise at least 10, at least 20, at least 50, at least 100, at least 1000, at least 5000, at least 10,000, at least 20,000 or up to 100,000 different peptides. Each population of antigen presenting cells may be modified with the entire library or a subset of the library. When using expression vectors, each population of antigen presenting cells to be incubated with a population of T cells may be modified such that each cell randomly expresses one candidate peptide of the plurality of candidate peptides. Indeed, in such embodiments the identity of the expressed candidate peptide for any selected doublet can be recovered by sequencing. When using peptides, if the identity of the peptides involved in any selected doublet is to be determined in the assay, then it may be advantageous for each coculture to include cells that present a single peptide. When using an expression vector library, each expression vector of the library can comprise, in addition to a sequence encoding for a respective peptide, one or more optional components selected from: a barcode sequence, a poly-A tail and / or a pair of priming sites. The pair of priming sites may be universal priming sites that have a sequence that is not expected to occur in the antigen presenting cells in the absence of the expression vector. At step 216, a population of T cells is obtained. The T cells may be unmodified and / or primary T cells. In other words, the T cells may be T cells that have been obtained from a patient and have not been subject to any genetic modification (Although the cells may have been subject to one or more expansion steps). The population of T cells may comprise CD8 T cells. In such embodiments, the selected MHC molecule used may be a class I molecule. The population of T cells may comprise CD4 T cells. In such embodiments, the selected MHC molecule used may be a class II molecule. The population of T cells may comprise CD8 T cells and CD4 T cells. In such embodiments, the selected MHC molecule used may be a class I molecule and / or a class II molecule (i.e. both a class I molecule and a class II molecule may be used in the assay, in the same or separate antigen presenting cells). Step 216 may comprise obtaining a sample comprising naive T cells. Step 216 may optionally further comprise subjecting the naive T cells to one or more expansion steps. For example, step 216 may comprise performing a polyclonal expansion of the naive T cells in the presence of CD3 / CD28 activators and IL2. CD3 / CD28 activators may be humanized CD3 and CD28 agonists such as those present in T Cell TransActTM.
[0118] At step 218, the population of antigen presenting cells obtained at steps 212 / 214 is incubated in the presence of the population of T cells obtained at step 216, and optionally a degradation ligand associated with the reporter fusion protein (if used). In embodiments that make use of a degradation ligand, the coculture (incubating step) can be performed in the presence of the degradation ligand from the start of the coculture. Alternatively, the degradation ligand can be included in the coculture during the incubating step, such as e.g. after a predetermined period of time in coculture. For example, the population of antigen presenting cells may be incubated in the presence of the population of T cells in the absence of the degradation ligand for a first predetermined period of time (such as e.g. between 2 and 24 hours, at least 4 hours, between 2 and 6 hours, between 2 and 24 hours, or between 12 and 24 hours), then the degradation ligand may be added and the cells cultured for a second predetermined period of time (such as e.g. between 1 and 4 hours, e.g. 3 hours). Alternatively, the population of antigen presenting cells may be incubated in the presence of the population of T cells and the degradation ligand for a predetermined period of time. Incubation in the presence of the degradation ligand can be achieved by adding the degradation ligand to the culture medium in which the cells are incubated. The present inventors have found the use of a coculture for a first predetermined period of time in the absence of the degradation ligand followed by coculture for a second predetermined period of time in the presence of the degradation ligand to be particularly beneficial for the sensitive detection of T- cell activation events. The coculture (incubating step) may be performed for a predetermined period of time (e.g. a first predetermined period of time, a predetermined period of time or a combined first and second predetermined periods of time) between 2 hours and 24 hours. For example, the incubating may be performed for at most 24 hours, at most 16 hours, at most 12 hours, at most 5 hours, or at most 3 hours, between 12 and 24 hours, about 16 hours, between 4 and 12 hours, about 12 hours, between 2 and 6 hours, about 3 hours, about 4 hours, or about 5 hours. The use of about 16 hours may be particularly convenient as it can be performed as an overnight culture. The use of between 12 and 24 hours time periods may be particularly useful when the antigens are expressed by the cells. The use of between 2 and 6 hours may be particularly useful when the cells are pulsed with antigen peptides. The population of antigen presenting cells and the population of T cells may be incubated using a 1 :10 to 10:1 ratio of cell numbers. In embodiments, approximately the same number of cells from the two populations (1 :1 ratio) is used. In other embodiments, a higher number of T cells compared to the number of antigen presenting cells is used, such as e.g. 2x to 10x as many T cells compared to antigen presenting cells. The use of higher number of T cells may improve sensitivity of detection. The coculture may be performed using between 1 million and 30 million, between 1 million and 20 million, or between 1 and 10 million antigen presenting cells. The number of cells used may depends on the time point used for selection (step 220) as well as the size of the library of antigens to be screened (i.e. the number of candidate antigens). Indeed, the later the selection step is performed the more likely an antigen presenting cell presenting a peptide recognised by a T cell will have undergone apoptosis, and the more copies of each antigen presenting cell presenting the same peptide will be needed to ensure that at least one doublet is recovered.
[0119] At step 220, one or more doublets of cells are selected from the cell culture obtained at step 218 (also referred to herein as “co-culture”), the selected doublets being positive for a marker included in the reporter fusion protein. As such, selecting the doublets enables isolation of an immunogenic peptide and a T cell expressing a TCR reactive to the peptide in the context of the selected HLA allele. This step may be performed using a cell sorting assay, such as e.g. FACS (fluorescence activated cell sorting), or magnetic cell sorting (MACS, e.g. using cell separation beads). At steps 222 / 224, nucleic acid material isolated from one or more of the selected doublets is sequenced, thereby identifying a candidate peptide associated with the antigen presenting cell in the doublet as an immunogenic peptide and / or a TCR expressed by the T cell in the doublet as a reactive TCR. This may comprise a step 222 of obtaining a single cell sequencing library in which nucleic acid material isolated from each doublet is associated with a respective barcode. The step of obtaining a single cell sequencing library can comprise capture of polyA containing nucleic acids (in the case of single cell RNA sequencing), for example by performing polyT based capture and reverse transcription. The step of obtaining a single cell sequencing library can comprise targeted amplification of the coding sequence of the peptide introduced into the antigen presenting cells, the coding sequence of the MHO molecule introduced in the antigen presenting cell, and / or the TCR sequence (or a part thereof) present in the T cell, as explained elsewhere. The single cell sequencing may be RNA sequencing or DNA sequencing depending on the implementation of step 222. The step of obtaining a sequencing library typically includes a single cell processing apparatus in which single cells (or in this case doublets of cells) are individually processed at least for the steps of nucleic acid isolation (e.g. cell lysis) and barcoding. Examples of single cell processing apparatus include the BD Rhapsody platform (for single cell RNA sequencing) and the Tapestri platform (for single cell DNA sequencing). At step 224, the library obtained at step 222 is sequenced, typically using any next generation sequencing method known in the art. This typically comprises obtaining sequence data comprising a plurality of sequencing reads. The step of sequencing the nucleic acid material may further comprise analysing the sequence data to identify a TCR sequence and / or a peptide sequence and / or a MHC sequence. This may comprise sorting the sequence data by doublet based on the single cell sequencing barcodes. This may further comprise, ad for a doublet to be analysed, aligning the associated sequencing reads to a suitable reference. The reference may include one or more MHC allele sequences, coding sequences for peptides in a library of candidate peptides, and / or TCR sequences. Further, when the sequence data is RNA sequence data, analysis of the sequence data may comprise analysing the sequence data to detect the presence or absence of a gene expression signature of T cell activation, wherein the presence of said signature further confirms that the doublet comprises an immunogenic peptide and T cell recognising said peptide. Note that the step of analysing sequence data is typically computer implemented. Indeed, the typically size and complexity of next generation sequencing data sets is such that analysis is beyond the capability of the human mind (e g. typical datasets comprise thousands of reads that must each be aligned to a reference sequence that is typically in the order of 3Gb for a human genome or at least thousands of kb for libraries of antigen coding sequences and TCR sequences).
[0120] At optional step 226, the results of one or more of the preceding steps are provided to a user, for example through a user interface.
[0121] The methods described herein find applications in any context in which it is desirable to identify and / or isolate one of a plurality of antigens and / or one of a plurality of T cells or corresponding TCRs that are capable of forming an immunogenic triplet in the context of a particular MHC molecule, as well as any context in which it is desirable to assess the effect of one or more perturbations on antigen recognition by T cells. For example, the methods described herein find applications in the context of designing and / or making an immunotherapy. An immunotherapy refers to a therapeutic or prophylactic compound or composition (including but not limited to compositions comprising cells or populations of cells, compositions comprising viral particles, compositions comprising a plurality of small and / or large molecules, compositions comprising nucleic acid material) that exploits or manipulates the immune system. An immunotherapy may comprise one or more antigens or coding sequences thereof, one or more T cell receptor or TCR- like molecules (e.g. TCR mimic antibodies), one or more T cell populations (including e.g. TCR-T cells, neoantigen specific tumour infiltrating lymphocytes, CAR-T cells), one or more antigen presenting cell populations, and / or one or more compounds that modulate an immune response by T cells (including compounds that modulate T cell activity - including but not limited to cytotoxicity - and / or health, and / or compounds that recruit T cells such as e.g. bi-specific T-cell engagers). The methods described herein find applications in the context of any disease or disorder which is related to or otherwise involves antigen recognition, including for example in cancer, infectious diseases and autoimmunity.
[0122] Also described herein is a method of screening a library comprising a plurality of candidate peptides to identify one or more immunogenic antigens, the method comprising: incubating a population of antigen presenting cells in the presence of a population of T cells, wherein the population of antigen presenting cells is a population of cells modified to express a reporter fusion protein as described herein; selecting one or more antigen presenting cells that express the reporter fusion protein or doublets of cells that comprise a T cell and an antigen presenting cell that expresses the reporter fusion protein using a selection method associated with the selection marker comprised in the reporter fusion protein, thereby isolating an immunogenic peptide or an immunogenic peptide and a T cell expressing a TCR reactive to the peptide. The cells presenting one or more of the plurality of candidate peptides may comprise respective constructs comprising coding sequences for the candidate peptides, and the method may comprise sequencing transcriptomic material isolated from one or more of the selected cells or doublets, thereby identifying a candidate peptide associated with the antigen presenting cell or antigen presenting cell in the doublet as an immunogenic peptide. Incubating a population of antigen presenting cells in the presence of a population of T cells may comprise incubating a plurality of subpopulations of antigen presenting cells in the presence of a population of T cells, each subpopulation comprising cells presenting a respective set of one or more of the plurality of candidate peptides on their cell surface in complex with a MHC molecule. The MHC molecule may be a MHC molecule encoded by a selected HLA allele. In such embodiments, the method may comprise identifying a candidate peptide or set of peptides as an immunogenic peptide or set of peptides based on the selection of doublets involving an antigen presenting cell of the respective subpopulation.
[0123] Also described herein is a method of identifying one or more TCRs that recognise an antigen, the method comprising: incubating a population of antigen presenting cells in the presence of a population of T cells, wherein the population of antigen presenting cells is a population of cells modified to express a reporter fusion protein as described herein, the population of antigen presenting cells comprising cells presenting the antigen on their cell surface in complex with a MHC molecule, optionally a MHC molecule encoded by a selected HLA allele that the antigen presenting cells have been modified to express; selecting one or more doublets of cells that comprise a T cell and an antigen presenting cell expressing the reporter fusion protein using a selection method associated with the selection marker comprised in the reporter fusion protein, thereby isolating a T cell expressing a TCR reactive to the antigen (optionally in the context of the selected HLA allele); and sequencing genetic and / or transcriptomic material from one or more cells of the selected doublets (e.g. by performing TCR sequencing), thereby identifying a TCR that recognises the antigen. The method may have any of the characteristics described herein in relation to a method of isolating an immunogenic peptide and / or a T cell receptor (TCR) reactive to a peptide.
[0124] Also described herein is a method of selecting one or more T cells that recognise one or more antigens from a sample comprising a plurality of T cell populations, the method comprising: incubating a population of antigen presenting cells in the presence of the sample comprising the plurality of population of T cells, wherein the population of antigen presenting cells is a population of cells modified to express a reporter fusion protein as described herein, the population of antigen presenting cells comprising cells presenting the one or more antigens on their cell surface, optionally in complex with a MHC molecule encoded by a selected HLA allele that the antigen presenting cells have been modified to express; selecting one or more doublets of cells that comprise a T cell and an antigen presenting cell that expresses the reporter fusion protein using a selection method associated with the selection marker comprised in the reporter fusion protein, thereby isolating a T cell expressing a TCR reactive to an antigen or the one or more antigens, optionally in the context of the selected HLA allele. The method may have any of the characteristics described herein in relation to a method of isolating an immunogenic peptide and / or a T cell receptor (TCR) reactive to a peptide.
[0125] Also described herein is a method of screening one or more compounds or compositions for an effect on T cell activation and / or antigen recognition and / or TCR mediated cytotoxicity, the method comprising: for each of the one or more compounds or compositions to be screened, incubating a population of antigen presenting cells in the presence of a population of T cells and the compound or composition to be screened, wherein the population of antigen presenting cells is a population of cells modified to express reporter fusion protein as described herein, the population of immortalised cells comprising cells presenting the antigen on their cell surface in complex with a MHC molecule, optionally a MHC molecule encoded by a selected HLA allele that the antigen presenting cells have been modified to express; selecting one or more antigen presenting cells that express the reporter fusion protein or doublets of cells that comprise a T cell and an antigen presenting cell that expresses the reporter fusion protein using a selection method associated with the selection marker comprised in the reporter fusion protein; and comparing the number of selected cells or doublets to a reference, thereby determining the effect of the compound or composition on T cell activation and / or antigen recognition. The method may have any of the characteristics described herein in relation to a method of isolating an immunogenic peptide and / or a T cell receptor (TCR) reactive to a peptide. The method may further comprise sequencing transcriptomics and / or genomic material from the selected doublets and comparing the sequence data to reference sequence data.
[0126] Also described herein is a method of measuring T cell activation and / or antigen recognition and / or TCR mediated cytotoxicity associated with a population of T cells (for example in the context of characterising a T cell based therapeutic, such as a CAR-T cell population, TIL population or TCR- T cell population), the method comprising: incubating a population of antigen presenting cells in the presence of the population of T cells, wherein the population of antigen presenting cells is a population of cells modified to express reporter fusion protein as described herein, the population of antigen presenting cells comprising cells presenting the antigen on their cell surface in complex with a MHC molecule, optionally a MHC molecule encoded by a selected HLA allele that the antigen presenting cells have been modified to express; selecting one or more antigen presenting cells that express the reporter fusion protein or doublets of cells that comprise a T cell and an antigen presenting cell that expresses the reporter fusion protein using a selection method associated with the selection marker comprised in the reporter fusion protein; and comparing the number of selected cells or doublets to a reference, thereby measuring T cell activation and / or antigen recognition and / or cytotoxicity associated with the T cell population. The method may have any of the characteristics described herein in relation to a method of isolating an immunogenic peptide and / or a T cell receptor (TCR) reactive to a peptide. The method may further comprise sequencing transcriptomics and / or genomic material from the selected doublets and comparing the sequence data to reference sequence data.
[0127] Also described herein is a method of obtaining data for training a machine learning model for prediction of immunogenicity, and a method of training a machine learning model for prediction of immunogenicity using said training data, the method comprising: incubating a population of antigen presenting cells in the presence of a population of T cells, optionally wherein the population of antigen presenting cells is a population of cells modified to express a selected HLA allele, the population of antigen presenting cells comprising cells presenting one or more of a plurality of candidate peptides on their cell surface in complex with a MHC molecule optionally encoded by the selected HLA allele, wherein said cells comprise respective constructs comprising coding sequences for the candidate peptides; selecting one or more doublets of cells that comprise a T cell and an antigen presenting cell that expresses the reporter fusion protein using a selection method associated with the selection marker comprised in the reporter fusion protein, thereby isolating an immunogenic peptide and a T cell expressing a TCR reactive to the peptide in the context of the selected HLA allele; and sequencing transcriptomic material isolated from one or more of the selected doublets, thereby identifying a candidate peptide associated with the antigen presenting cell in the doublet as an immunogenic peptide and optionally a TCR that recognises said immunogenic peptide. The training data may comprise, for each of the plurality of candidate peptides, an indication of whether the peptide was identified as an immunogenic peptide. The training data may further comprise an indication of the selected HLA allele and / or the TCR that was identified as recognising an immunogenic peptide. The machine learning model may be a machine learning model trained to predict whether a peptide is likely to be an immunogenic peptide, optionally in the context of a selected HLA allele and / or in relation to an identified TCR. The method may further comprise training the machine learning model. The machine learning model may be a machine learning model as described in PCT / EP2024 / 057046. The machine learning model may be a machine learning model as described in Lu et al. 2021 (PTMnet), Moris et al., 2020 (Imrex), or O’Donnell et al., 2020 (MHCflurry), or any other model for the prediction of antigen presentation and / or immunogenicity. The method may further comprise training said model using said training data. The method may have any of the characteristics described herein in relation to a method of isolating an immunogenic peptide and / or a T cell receptor (TOR) reactive to a peptide.
[0128] The above methods find applications in the context of designing immunotherapies, particularly immunotherapies that use peptides or sequences encoding peptides to generate or promote an immune response, and therapies that use reactive T cells or reactive TOR constructs as therapeutic agents. Indeed, peptides that are verified to be immunogenic (particularly in the context of a specific set of candidate MHC molecules and / or candidate TCR molecules, such as e.g. based on the TCR repertoire and / or MHC alleles identified to be present in a sample or patient) are more promising candidates for inclusion in the immunotherapy. In particular, the above methods may be used to provide cancer immunotherapies that target cancer-specific antigens (also referred to herein as “cancer neoantigens”, or simply “neoantigens”). As the skilled person understands, a cancer-specific antigen may be truly specific to cancer cells (in the sense that it is only expressed by the genome of cancer cells), or may be practically specific to cancer cells (in the sense that it is expressed by cancer cells at a significantly higher level than by normal cells). The cancer neoantigens may be clonal neoantigens. Thus, also described herein are methods of providing an immunotherapy for a subject, the method comprising identifying and optionally producing one or more peptides that comprise cancer neoantigens determined to be immunogenic using a method as described herein. Further, also described herein are methods of providing an immunotherapy for a subject, the method comprising identifying one or more TCR sequences or T cells that are determined to be reactive using a method as described herein, and producing a therapy using said cells or TCR sequences. Such a therapy may be, e.g. expanded autologous or heterologous T cells, such as autologous tumour infiltrated lymphocytes (TILs) (e.g. where reactive T cells are isolated using a method as described herein, and expanded before being administered to a patient), modified T cells such as TCR-T and CAR-T cells (e.g. where cells are modified to express a TCR or part thereof identified using a method as described herein), or therapeutic compound that include a TCR or part thereof identified using a method as described herein, such as a bispecific T cell engages (BiTE) or a TCR mimetic antibody.
[0129] Examples of method of providing immunotherapies will be described by reference to Figure 3B.
[0130] Figure 3B illustrates schematically an exemplary method of providing an immunotherapy. At optional step 308, one or more samples comprising tumour genetic material and one or more germline samples are obtained from a subject or a plurality of subjects. The subjects may be subjects that have been diagnosed as having cancer. The subject may be (but does not need to be) the same subject for which the immunotherapy is provided. Step 308 may be omitted and the method may instead start from a list of candidate neoantigens or from previously acquired sequence data from which candidate neoantigens are identified. At step 310, a list of candidate antigens is obtained using methods known in the art, for example as described in WO 2022 / 207925 and WO 2016 / 16174085, and others. In the illustrated embodiment the antigens are neoantigens, btu the methods described herein are in principle applicable to any other type of antigens. The neoantigens may be clonal neoantigens. Methods to identify clonal neoantigens are known in the art and include the methods described in WO 2022 / 207925, WO 2016 / 16174085, and McGranahan et al. (2016). The clonal neoantigens may in particular be identified using a method as described in WO 2022 / 207925. The list may comprise a single neoantigen, or a plurality of neoantigens. Preferably, the list comprises a plurality of neoantigens. At optional step 312, one or more candidate peptides are identified for each of the candidate neoantigens. For example, a plurality of peptides may be designed for at least one of the candidate clonal neoantigens, which differ in their lengths and / or the location of a sequence variation that characterises the neoantigen compared to the corresponding germline peptide. Alternatively, one or more sequences encoding the candidate neoantigens may be identified.
[0131] At step 314, one or more populations of antigen presenting cells displaying the one or more peptides are obtained, for example by pulsing the cells with the candidate peptides or by introducing sequences encoding the candidate neoantigens in the cells (e g. using one or more tandem minigenes). The population of cells may express a MHC molecule encoded by a selected HLA allele determined to be present in a tumour of a subject to be treated, such as e.g. the subject from which samples have been obtained at step 306. At step 316, one or more T cell populations are obtained, for example from a subject to be treated (e.g. TILs or blood derived T cells), or from healthy subjects. At step 318, the population(s) of antigen presenting cells obtained at step 314 are incubated with the one or more T cell populations obtained at step 316, and optionally with a degradation ligand, as explained by reference to Figure 2. At step 320 antigen presenting cells or doublets of cells are selected as explained by reference to Figure 2. At optional steps 322, 324, transcriptomic material is isolated from doublets and sequence data is analysed as explained by reference to Figure 2.
[0132] At optional step 326B, one or more of the peptides are selected for production based on at least some of the results of step 324. For examples, peptides identified as immunogenic peptides or sequences encoding such peptides may be selected for manufacture. Peptides with selected sequences may be obtained using any method known in the art but they are preferably obtained using chemical synthesis. Methods for obtaining sequences that encode peptides of interest are known. For example, tandem minigenes may be obtained which encode the selected one or more peptide. At step 328B, an immunotherapy may be produced using at least some of the one or more peptides or sequences encoding said peptides produced at step 326B. The immunotherapy may comprise the one or more peptides (e.g. in the case of an immunogenic composition such as a synthetic long peptide vaccine), sequences encoding said peptides (e.g. in the case of a DNA or RNA vaccine) or may comprise molecules or cells that have been obtained using the selected peptides (e.g. in the case of therapeutic antibodies that selectively bind the candidate peptides, or immune cells that specifically recognise the candidate peptides). For example, the immunotherapy may comprise cells that have been obtained using the selected peptides. Methods of producing an immunotherapy comprising cells that have been obtained using neoantigen peptides are known in the art, for example as described in WO2022 / 269250, WO 2022 / 207925, WO 2016 / 16174085, and McGranahan et al. (2016). The immunotherapy may comprise cells that have been obtained by expansion of a T cell population that is enriched forT cells selected using a method as described herein. Such a population may be obtained by selective expansion of a population comprising T cells in the presence of an antigen presenting cell and one or more of the peptides selected at step 326B. Instead or in addition to this, such a population may be obtained by expansion of a population comprising T cells that have been isolated using a method as described herein. For example, at optional step 326A, T cells from the doublets selected at step 320 may be isolated. At step 328A, the selected T cells are used to produce an immunotherapy, for example by expanding the selected T cells.
[0133] At optional step 326C, the sequence of one or more reactive TCRs identified at step 324 are used to design an immunotherapy comprising a population of T cells modified to express the one or more reactive TCRs or a part thereof, such as e.g. a CAR-T cell population. At optional step 328B, an immunotherapy comprising the one or more reactive TCRs or part thereof is obtained.
[0134] At optional step 330, any of the immunotherapies obtained in the preceding steps may be administered to a subject. This may be the subject from which the samples used to identify the neoantigens have been obtained, or another subject.
[0135] Note that in principle the method of Figure 3B is applicable to any antigens, and is not limited to neoantigens. Thus, also described herein is a method comprising any one or more of steps 310 to 330 of Figure 3B, in which the list of candidate antigens are antigens associated with a pathogen, allergens or autoantigens.
[0136] A cancer immunotherapy refers to a therapeutic approach comprising administration of an immunogenic composition (e.g. a vaccine), a composition comprising immune cells, or an immunoactive drug, such as e.g. a therapeutic antibody, to a subject who has been identified as having or being likely to have or develop cancer. The term “immunotherapy” may also refer to the therapeutic compositions themselves. A cancer immunotherapy can target a neoantigen or cancer associated antigen. For example, an immunogenic composition or vaccine may comprise a neoantigen, neoantigen presenting cell or material necessary for the expression of the neoantigen. As another example, a composition comprising immune cells may comprise T and / or B cells that recognise a neoantigen. The immune cells may be isolated from tumours or other tissues (including but not limited to lymph node, blood or ascites), expanded ex vivo or in vitro and re-administered to a subject (a process referred to as “adoptive cell therapy”). The expansion step may use a peptide identified at step 326B, or may include cells isolated at step 326A. Instead or in addition to this, T cells can be isolated from a subject and engineered to target a neoantigen (e.g. by insertion of a chimeric antigen receptor that binds to the neoantigen) at step 326C, and re-administered to the subject. As another example, a therapeutic antibody may be an antibody which recognises a neoantigen, produced at step 328B using one or more peptides selected at step 326B. For example, antibody libraries may be tested for binding to a peptide selected at step 326B. One skilled in the art will appreciate that if the neoantigen is a cell surface antigen, an antibody as referred to herein will recognise the neoantigen. Where the neoantigen is an intracellular antigen, the antibody will recognise the neoantigen peptide-MHC complex. As referred to herein, an antibody which "recognises" a neoantigen encompasses both of these possibilities. Further, an immunotherapy may target a plurality of neoantigens. For example, an immunogenic composition may comprise a plurality of neoantigens, cells presenting a plurality of neoantigens or the material necessary for the expression of the plurality of neoantigens. As another example, a composition may comprise immune cells that recognise a plurality of neoantigens. Similarly, a composition may comprise a plurality of immune cells that recognise the same neoantigen. As another example, a composition may comprise a plurality of therapeutic antibodies that recognise a plurality of neoantigens. Similarly, a composition may comprise a plurality of therapeutic antibodies that recognise the same neoantigen. References to "an immune cell" are intended to encompass cells of the immune system, for example T cells, NK cells, NKT cells, B cells and dendritic cells. In a preferred embodiment, an immune cell for use as an immunotherapy is a T cell. An immune cell that recognises a neoantigen may be an engineered T cell. A neoantigen specific T cell may express a chimeric antigen receptor (CAR) or a T cell receptor (TCR) which specifically binds a neoantigen, or an affinity-enhanced T cell receptor (TCR) which specifically binds a neoantigen, and which as been derived from a TCR sequence identified using a method as described herein (Step 326C). For example, the T cell may express a chimeric antigen receptor (CAR) or a T cell receptor (TCR) which specifically binds to a neoantigen (for example an affinity enhanced T cell receptor (TCR) which specifically binds to a neo-antigen or a neo-antigen peptide). Alternatively, a population of immune cells that recognise a neoantigen may be a population of T cell isolated from a subject with a tumour. For example, the T cell population may be generated from T cells in a sample isolated from the subject, such as e.g. a tumour sample, a peripheral blood sample or a sample from other tissues of the subject. The T cell population may be generated from a sample from the tumour in which the neoantigen is identified. In other words, the T cell population may be isolated from a sample derived from the tumour of a patient to be treated, where the neoantigen was also identified from a sample from said tumour. The T cell population may comprise tumour infiltrating lymphocytes (TIL).
[0137] The term "Antibody" (Ab) includes monoclonal antibodies, polyclonal antibodies, multispecific antibodies (e.g., bispecific antibodies), and antibody fragments that exhibit the desired biological activity. The term "immunoglobulin" (Ig) may be used interchangeably with "antibody". Once a suitable neoantigen has been identified, for example by a method according to the disclosure, methods known in the art can be used to generate an antibody.
[0138] A composition as described herein may be a pharmaceutical composition which additionally comprises a pharmaceutically acceptable carrier, diluent or excipient. The pharmaceutical composition may optionally comprise one or more further pharmaceutically active polypeptides and / or compounds. Such a formulation may, for example, be in a form suitable for intravenous infusion.
[0139] An “immunogenic composition” is a composition that is capable of inducing an immune response in a subject. The term is used interchangeably with the term “vaccine”. The immunogenic composition or vaccine described herein may lead to generation of an immune response in the subject. An "immune response" which may be generated may be humoral and / or cell-mediated immunity, for example the stimulation of antibody production, or the stimulation of cytotoxic or killer cells, which may recognise and destroy (or otherwise eliminate) cells expressing antigens corresponding to the antigens in the vaccine on their surface. The immunogenic composition may comprise one or more antigens, or the material necessary for the expression of one or more antigens. In addition, an antigen may be delivered in the form of a cell, such as an antigen presenting cell, for example a dendritic cell. The antigen presenting cell such as a dendritic cell may be pulsed or loaded with the antigen or antigen peptide or genetically modified (via DNA or RNA transfer) to express one, two or more antigens or antigen peptides, for example 2, 3, 4, 5, 6, 7, 8, 9 or 10 antigens or antigen peptides, where at least one of the antigen peptides has been identified using a method as described herein. Methods of preparing dendritic cell immunogenic compositions or vaccines are known in the art.
[0140] The immunotherapies described herein may be used in the treatment and / or prevention of cancer (including cancer recurrence). For example, immunotherapeutic compositions may be used to produce or potentiate an immune response to an existing cancer, or to prime a subject’s immune system to recognise a cancer neoantigen. Thus, the disclosure also provides a method of treating and / or preventing cancer in a subject comprising administering an immunotherapeutic composition as described herein to the subject.
[0141] As used herein "treatment" refers to reducing, alleviating or eliminating one or more symptoms of the disease which is being treated, relative to the symptoms prior to treatment. "Prevention" (or prophylaxis) refers to delaying or preventing the onset of the symptoms of the disease. Prevention may be absolute (such that no disease occurs) or may be effective only in some individuals or for a limited amount of time.
[0142] Suitably, in any embodiment of any aspect described herein, the cancer may be ovarian cancer, breast cancer, endometrial cancer, kidney cancer (renal cell), lung cancer (small cell, non-small cell and mesothelioma), bladder cancer, gastric cancer, oesophagal cancer, colorectal cancer, cervical cancer, endometrial cancer, brain cancer (gliomas, astrocytomas, glioblastomas), melanoma, merkel cell carcinoma, clear cell renal cell carcinoma (ccRCC), lymphoma, small bowel cancers (duodenal and jejunal), leukemia, pancreatic cancer, hepatobiliary tumours, germ cell cancers, prostate cancer, head and neck cancers, thyroid cancer and sarcomas. For example, the cancer may be lung cancer, such as lung adenocarcinoma or lung squamous-cell carcinoma. As another example, the cancer may be melanoma. The cancer may be bladder cancer. The cancer may be head and neck cancer. In embodiments, the cancer may be selected from melanoma, merkel cell carcinoma, renal cancer, non-small cell lung cancer (NSCLC), urothelial carcinoma of the bladder (BLAC) and head and neck squamous cell carcinoma (HNSC) and microsatellite instability (MSI)-high cancers. In some embodiments, the cancer is non-small cell lung cancer (NSCLC). In any embodiment of any aspect, the subject may be human.
[0143] Treatment using the compositions and methods of the present disclosure may also encompass targeting circulating tumour cells and / or metastases derived from the tumour. Treatment according to the present disclosure targeting one or more neoantigens, preferably clonal neoantigens, may help prevent the evolution of therapy resistant tumour cells which may occur with standard approaches such as chemotherapy, radiotherapy, or non-specific immunotherapy. The methods and uses for treating cancer described herein may be performed in combination with additional cancer therapies. In particular, the immunotherapies (including but not limited to T cell compositions) described herein may be administered in combination with immune checkpoint intervention, co-stimulatory antibodies, chemotherapy and / or radiotherapy, targeted therapy or monoclonal antibody therapy. 'In combination' may refer to administration of the additional therapy before, at the same time as or after administration of the immunotherapy (e g. T cell composition) as described herein.
[0144] Also described herein is a method of treating a subject that has been diagnosed as having cancer, the method comprising administering an immunotherapy that has been designed or provided using the methods described herein, or a composition as described herein.
[0145] Further, the present disclosure provides methods for determining the effect of one or more test conditions (e.g. exposure to one or more compounds or compositions, co-culture with one or more cell populations, exposure to one or more physico-chemical perturbations, etc) on apoptosis of a target cell population, or gamma secretase activity in a target cell population. Gamma secretase activity has bene implicated in the cleavage of amyloid precursor protein to generate extracellular Ap peptides. Extracellular Ap peptides are the main components of extracellular amyloid plaques, a hallmark of Alzheimer’s disease (AD). An illustrative method will be described by reference to Figure 4. In such embodiments, the protease recognition sequence is a recognition sequence of an apoptosis associated protease, such as caspase-3 / 7 or caspase-8 / 10, or a recognition sequence of gamma secretase. Methods described by reference to Figure 4 find use for example in the context of high throughput drug screening assays, for example where the target cells are cancer cells or central nervous system cells. At step 410, a population of target cells is obtained which is a population of cells modified to express a reporter fusion protein as described herein. This may comprise introducing an expression vector into the cells, the expression vector comprising a nucleic acid sequence encoding the receptor fusion protein. At step 412, the cells are cultured in the presence off the one or more test conditions, and optionally the degradation ligand associated with the reporter fusion protein (if any). The target cells may be e.g. cancer cells, or central nervous system cells. At step 414, cells that express the reporter fusion protein are selected or selectively quantified using a selection method associated with the selection marker comprised in the reporter fusion protein. For example, positive cells may be counted by flow cytometry, quantified by FACS, selected and / or quantified by magnetic bead based cell sorting, etc. Selected cells may optionally be further characterised using one or more known characterisation methods such as e.g. transcriptomic characterisation, genetic characterisation, morphologic characterisation, proteomic characterisation, etc. At step 416, the results of step 414 are compared between a plurality of test conditions or between one or more test conditions and one or more control values, thereby determining the effect of the respective test conditions on apoptosis of the target cells.
[0146] Systems
[0147] Figure 5 shows a system for use in methods according to embodiments of the present disclosure. The system comprises a computing device 50, which comprises a processor 501 and computer readable memory 502. In the embodiment shown, the computing device 50 also comprises a user interface 503, which is illustrated as a screen but may include any other means of conveying information to a user such as e.g. through audible or visual signals. The computing device 50 is communicably connected, through a wired or wireless connection such as e.g, through a network (not shown), to sequence data acquisition means 54, such as a sequencing machine, and / or to one or more databases 53 storing sequence data. The one or more databases may additionally store other types of information that may be used by the computing device 50, such as e.g. reference sequences, parameters, etc. The computing device is configured to implement any computer-implemented method step described herein, such as e.g. analysing sequence data, identifying a plurality of peptides, etc. The connection between the computing device 50 and the sequence data acquisition means 54 may be direct or indirect (such as e.g. through a remote computer). The sequence data acquisition means 54 are configured to acquire sequence data from nucleic acid samples, for example RNA samples or DNA samples derived therefrom (e.g. for the purpose of analysing transcriptomic material isolated from doublets of cells) and also optionally genomic DNA samples (e.g. for the purpose of identifying candidate peptides) extracted from cells and / or tissue samples. The sequence data acquisition means 54 preferably comprises a next generation sequencer. The sequence data acquisition means 54 may be in direct or indirect connection with one or more databases 55, on which sequence data (raw or partially processed, such as e.g. raw reads, aligned reads, identified variants, etc.) may be stored. The system as illustrated further comprises a cell culture system 52, such as e.g. an incubator, in which cells can be maintained in conditions compatible with cell viability, as well as a cell sorting device 53. The cell sorting device may be flow cytometer-based device such as e.g. a FACS machine. The cell sorting device 53 may be used to separate doublets of cells as described herein, from a cell culture product after incubation in the cell culture system 52. The doublets may then be subject downstream processing to recover the T cells (e.g. for T cell population enrichment) or for sequencing analysis (to identify any or all of the antigen, MHC molecule and TCR involved in the interaction between the cells in the doublet). When the doublets are prepared for sequencing analysis, any method known for preparation of libraries for single cell sequencing may be used. The single cell sequencing may be single cell RNA sequencing or single cell DNA sequencing. Doublets separated as described herein may each be recovered in a well of a multi-well plate. For example, doublets may be sorted into 384-well cell capture plates. Alternatively, the doublets may be processed using a single cell processing system such as BD Rhapsody, which captures individual cells (or here doublets) into respective partitions (microwells) within a cartridge. The partitions are loaded with magnetic beads which allow retrieval of the mRNA material in each well after cell lysis, and cDNA synthesis of the retrieved material. The wells / microwells may comprise a buffer or lysis solution. The wells / microwells may comprise one or more reagents for sequencing library preparation, or the RNA material obtained from the lysed cells may be barcoded, removed from the wells and processed for sequencing library preparation using the reagents for sequencing library preparation. These may include poly(T) reverse transcription primers for single-cell RNA- seq, and / or one or more primers for selective amplification of target sequences (e.g. primers for selective amplification of the MHC construct and / or primers for selective amplification of the antigen construct and / or primers for selective amplification of TCR sequences). Specific amplification of TCR sequences encompasses specific amplification of parts of TCR sequences, such as e.g. the part coding for the CDR3 region of TCRs. For example, specific amplification and sequencing of TCRs may be performed by preparing a sequencing library including a multiplex PCR step using a multiplex pool of forward PCR primers complementary to all the possible V segments and a pool of reverse primers complementary to the constant region of the alpha and beta TCR chains. As another example, specific amplification of TCR sequences may comprise amplification of CDR3 sequences, for example using the BD Rhapsody™ VDJ CDR3 Protocol. Alternatively, the doublets may be processed using a single cell analysis instrument comprising cell encapsulation, such as the Tapestri platform. Doublets may be encapsulated in a lysis solution, then droplets may be combined with barcoding beads and a reagent mixture for library preparation, such as e.g. primers for selective amplification of the MHC and / or antigen coding sequence introduced in the cells, and / or TCR genomic region in the T cell. In other words, doublets may be encapsulated and subject to targeted DNA single cell sequencing. Thus, the sequence data acquisition means 54 may include a single cell processing apparatus such as a Tapestri single cell analysis instrument or a BD Rhapsody single cell analysis system, and a sequencing machine. Depending on the single cell processing apparatus used, the library for sequencing may represent DNA or RNA information (i.e. the system as a whole may perform single cell (or in this case doublet) RNA or DNA sequencing).
[0148] The following is presented by way of example and is not to be construed as a limitation to the scope of the claims.
[0149] EXAMPLES
[0150] Example 1
[0151] Introduction
[0152] As explained above, identifying antigens that are immunogenic, and their cognate T cell receptors is still extremely challenging. To overcome this challenge, the present inventors have developed a reporter system where a target cell becomes fluorescent when it is specifically targeted by a reactive T-cell (Figures 1A, 1 B). Fluorescent cells can then be isolated by FACS as doublets (target cell and its specifically interacting T-cell) and the neoantigen and TCR alpha-beta pairs identified by single cell sequencing. The reporter system demonstrated in the present examples is based on a GFP chimaera where the GFP sequence is fused in-frame to the FKBP12F36V degradation tagging sequence, with the two entities spaced with the granzyme B cleavage sequence (Figure 1 B, top). Cells transduced with this construct strongly express GFP. On the addition of the dTAG molecule, which is a PROTAC-based heterobifunctional degrader that binds the FKBP12F36V sequence, the protein is brought to the CRBN proteolysis system and rapidly degraded (Figure 2C). The inventors have shown that K562 cells transduced with this construct strongly express GFP and become GFP negative after 4 hrs of dTAG-13 treatment, remaining so for at least 3 days. When these cells are pulsed with immunogenic CMV peptide and co-cultured with CMV specific T-cells, T-cells engage with K562 target cells, secrete Granzyme B, which cleaves the Granzyme B peptide sequence, releasing GFP from the FKBP12F36V tag so it is no longer degraded. Through proof-of-principle experiments the inventors showed the system to be highly sensitive and specific (Figures 8-10).
[0153] Since different proteases recognise different peptide sequences (eg DEVD (SEQ ID NO: 3) sequence is recognised by Caspase-3 / 7, ELQTD (SEQ ID NO: 4) is recognised by caspase-8 / 10), this method is easily adapted to detect activity of different proteases of interest. For instance, a construct encoding GFP-DEVD- FKBP12F36V could be utilised as a reporter of apoptosis, useful for high throughput drug screening assays in cancer cells for use by biopharma.
[0154] Methods
[0155] Target Cell lines. K562 retrovirally expressing HLA-A*02:01 (K562-A2) were acquired from Sine Hardrup lab (DTU, Denmark). LAMP1 was knocked out via CRISPRCas9 RNP using sgRNA sequence: ACAACGTGAGCGGCACCAAC (IDT technologies, SEQ ID NO: 1 1). LAMP1 KO cells were FACS sorted for CD107a negative cells and subsequently transduced with CD80 and CD86 retrovirus supernatant purchased from Mirjam Heemskerk (LUMC, The Netherlands). Single cell clones were generated by FACS sorting K562-A2 for HLA-A2HI, CD80HI, CD86HIand CD107Ane9and the chosen clone was named KA2_DECOD. KA2_DECOD was lentivirally transduced with each printed mutant (PM) library pool at MOI 4 unless otherwise stated (Table 1) and selected for 10 days by 1 ug / ml puromycin. Printed mutant library pools refer to libraries of sequences of a predetermined length (referred to as “tiles”) that together cover 600,000 different mutations and represent a high diversity neoantigen library including but not limited to ~20k common cancer- associated somatic nucleotide variants derived from the PCAWG database. The term “tile” is also used to refer to the coding sequence of specific antigens of the same predetermined length (e.g. CMV tile refers to a sequence of a predetermined length encoding a CMV antigen). The final cell line was named according to the library expressed, e.g KA2_DECOD_PM7. All K562 cell lines were cultured in RPMI (GIBCO cat:11875), 10% FCS (Sigma-Aldrich cat:F9665), 1 % Pen / strep (Gibco cat:15140) (RPMI10), and media was replenished every 2-3 days.
[0156] KA2_DECOD cell lines were lentivirally transduced with CleavER-GFP reporter (SEQ ID NO:12, encoding amino acid sequence SEQ ID NO: 10). On Day 7 after transduction, KA2_DECOD were FACS sorted for GFPhi, HLA-A2hi,CD80HI, CD86HIand CD107Ane9 and named KA2_CleavER.
[0157] Table 1. Printed mutant library tiles.
[0158] Virus production. 3x10® HEK293T cells were seeded on 10 cm tissue culture dishes and grown in DMEM (Merck cat:D6429) supplemented with 10% FCS and 1 % Pen / strep for 24 hours, until reaching 50% confluency. Cells were then transfected using the GeneJuice® Transfection Reagent (Merck cat:70967-5): 4 pg of lentiviral target plasmid was mixed with 4 pg of psPAX2 (packaging vector), 2 pg of VSV-G (envelope vector), 30 pl of GeneJuice and 470 pl of Opti-MEM reduced serum medium (Thermo Fisher cat:31985062), incubated at room temperature (RT) for 15 minutes and added to HEK293T cell culture. After 24 hours, the medium was changed to RPMI (supplemented with 10% FCS and 1 % Pen / strep). At 48hours the virus-containing medium from HEK293T dishes was collected, filtered through 0.45 pm nitrocellulose membrane and stored at - 80°C. The following constructs were produced:
[0159] CMVA2_TILE: vector=modified Pix302; promoter=CMV; selection=EF1 a-puromycin; cloning by Gibson assembly; printed mutant libraries 1-7: vector=modified Pix302; promoter=CMV; selection=EF1a- puromycin; cloning by Gibson assembly;
[0160] GFP-FKB12-CleavER: vector=modified Pix302; promoter=CMV; selection=GFP; cloning by Gibson assembly.
[0161] Viral transductions. For viral transduction, 0.5 x 106of target cells were resuspended in 1 ml of virus-containing medium and transferred to a well of 24-well plate. Polybrene (Santa Cruz, cat:L1322) was then added at 8 pg / ml, after which the plate was centrifuged at 2,500 RPM for 1 .5 hours at 37°C and returned to a CO2incubator. Following overnight incubation, cells were centrifuged briefly and medium was replaced with fresh RPMI-1640 (10% FCS). In constructs containing a puromycin resistance gene as a selection marker, 2 pg / ml of puromycin was added to the transduced cells 48 hours after initial transduction and maintained on puromycin for 10 days.
[0162] CMV-specific T-cell line. CMV-specific T-cells were generated from HLA-A*02:01 + healthy donor PBMCs (BiolVT). In short: autologous dendritic cells (DCs) were loaded with 100nM CMV-derived peptide NLVPMVATV (SEQ ID NO: 5, referred to as “NLV”) presented in HLA-A*02:01 , and subsequently co-cultured with freshly isolated CD8 T-cells for 14 days. T-cells were restimulated with peptide loaded DCs and cultured for a further 14 days before cryopreservation. T-cell media (TCM) was TEXMACS (Miltenyi Biotec cat:130-097-196), supplemented with 5% human serum (Sigma-aldrich, H3667), 1% Pen / Strep (Gibco cat:15140) and 100IU / ml IL2 (Proleukin, Clinigen healthcare Ltd). Media was replenished every 2-3 days. The frequency of CMV reactive T-cells present in the culture was determined by CMV-specific tetramer binding.
[0163] Library-specific T-cell lines - Naive CD8 T-cell isolation. Following an adapted Miltenyi Biotec protocol, untouched Naive CD8 T cells were isolated from healthy donor PBMCs using a combination of Naive Pan T cell isolation kit (Miltenyi Biotech™ cat: 130-097-095) and CD4 microbead kit (Miltenyi Biotech™ cat: 130-045-101 ). Manufacturers instructions were followed for Naive pan T cell isolation and directly after incubation with Biotin-Microbeads, CD4 microbeads were added at 20pl / 1x107cells and incubated for a further 10mins in the fridge before separation on an LS column (Miltenyi Biotech™ cat: 131 -042-401) as per instructions. Cold Separation Buffer was used formulated of PBS (Gibco cat: 14190-094), 2% FCS (Sigma-Aldrich cat:F9665), 2mM EDTA (Invitrogen cat: 15575-038) and 1 % pen / strep (Gibco cat:15140).
[0164] Library-specific T-cell lines - Naive CDS T cell Polyclonal expansion using TransAct™. Freshly isolated, untouched Naive CD8 T cells were subsequently stimulated with TransAct™ (Miltenyi Biotech™ cat: 130- 128-758) at 10ul / ml. T-cells were plated at 1x106 / ml in Naive T-with cell media TEXMACS (Miltenyi Biotec cat:130-097-196), supplemented with 5% human serum (Sigma-aldrich cat:H3667), 1 % Pen / Strep (Gibco cat:15140), 100IU / ml IL2 (Proleukin, Clinigen healthcare Ltd), 10ng / ml IL7 (CellGenix cat: 1010-050), 10ng / ml IL15 (CellGenix cat:1413-050) and on Day 0 only 50ng / ml IL21 (CellGenix cat: 1419-050). Media was replenished every 2-3 days and on Day 14 polyclonally expanded T-cells were cryopreserved and 2x106were prepped for RNA extraction. Phenotype of the T-cells were determined using TBNK + CD335-AF700 staining and a memory T-cell panel.
[0165] Library-specific T-cell lines - Library-specific T-cell expansion: KA2_DECOD cell lines lentivirally transduced with printed mutant library 7 (PM7) were resuspended at 10x106 / ml in RPMI10 and treated with 25ug / ml mitomycin C (cat:M5353, Sigma-Aldrich) for 30mins at 37°C. Cells were washed by centrifugation (450g 10mins) 4 times by diluting the cells 10x in RPMI. Unless otherwise stated, 20x106mitomycin treated KA2_DECOD_PM7 cells were resuspended in Naive T-cell media and cocultured with 20x106Naive polyclonally expanded CD8 T cells at a 1 :1 E:S ratio. Media was replenished every 2-3 days and on Day 14 cells were harvested and cryopreservation.
[0166] CMV expression. CMV peptide expressing KA2_CleavER cells were generated by loading with 100nM CMV-derived peptide NLVPMVATV (SEQ ID NO: 5) for 1 hr at 37°C in RPMI10 and excess peptide was removed by washing twice in RPM110 by centrifugation (450g 5mins). Transient endogenous CMV antigen expression was achieved by transfecting KA2_CleavER with 1 ug of tandem mini gene (TMG) encoding CMV 29mer amino acid sequence (pp65484-513) linked to CFP via a 2A sequence (referred to herein as TMG_NLV and provided as SEQ ID NO: 6). Stable endogenous CMV antigen expression was achieved using a lentiviral expression system, CMV 29mer amino acid sequence (pp65484-513) was ordered as a dsDNA TILE (Twist Bioscience) (sequence with vector overhang provided as SEQ ID NO: 13, sequence without vector overhang provided as SEQ ID NO: 14) and cloned via Gibson assembly into a modified Plx302-Puromycin vector. Lentivirus was produced as described and transduced into KA2_CleavER and transduced cells were selected using puromycin.
[0167] CleavER sensitivity assay. KA2_CleavER cells expressing CMV antigen were spiked into empty KA2_CleavER cells at described frequencies. Similarly, CMV-specific T-cells were spiked into purified CD8 T-cells derived from the same healthy donor. Spiked Targets and spiked T-cells were then combined at a 1 :1 E:S ratio and co-cultured for either 4 hours or 16hrs at 37°C in TEXMACS media (Miltenyi Biotec cat: 130-097-196) supplemented with 5% human serum (Sigma-aldrich, H3667) and 1 % pen / strep (Gibco cat:15140) (TCM) in the presence of anti CD107a-APC. Following incubation, dtAG-13 (cat:HY-1 14421 , Cambridge Bioscience) was added directly to the cultures at a final concentration 2uM and cultured for a further 3 hours at 37°C. Cell cultures were then washed once in Cold Separation Buffer and stained for surface markers CD8b-PE and CD86- BV786 or CD3-BV786 and CD80-BV421 with CD137-PeCy7. In the absence of CD3 or CD8 stain T-cells were labelled with cell trace violet prior to co-culture with Target cells. Stained cells were then measured by flow cytometry and analysed by Flow-jo software.
[0168] CleaVER Neoantigen screens. 10-20x10® PM7 library expanded T-cells were thawed and rested for 4 hours at RT in TCM at 1x106 / ml. PM7 expanded T-cells were then cocultured with KA2_CleavER_PM7 cells expressing the corresponding Printed mutant library a 1 :1 E:S ratio for 16hrs at 37°C in the presence of anti-CD107a-APC. Following incubation, DTAG-13 (HY-1 14421 , Cambridge Bioscience) was added directly to the cultures at a final concentration 2uM and cultured for a further 3 hours at 37°C. Cell cultures were then washed once in Cold Separation Buffer and stained for surface markers CD3-BV786, CD80-BV421 and CD137-PeCy7 for 30mins at 5x106 / ml. Cells were then washed and resuspended at 5x106 / ml in Cold Separation buffer and sorted by FACS.
[0169] FACS Sort gating strategy. Dead cells were removed by 7AAD staining which was added 5 mins prior to FACS sorting. KA2_Cleaver cells were gated on CD80 expression and GFP+ and GFP- cells were sorted into cold collection tubes containing RPMI10. Gates were determined using DTAG-13 treated KA2_CleavER_PM7 alone. T-cells were gated on CD3 expression and CD107a+CD137+ reactive T-cells and CD107a-CD137- non-reactive cells were sorted into cold collection tubes containing TCM. scRNA sequencing. 10-20,000 of each sorted cell population were prepped following BD Rhapsody ™ workflows. Indexed libraries were then sequenced via llumina Next Seq 2000 using a P1-300 cartridge system with a minimum of 2000 reads per cell for TILE identification.
[0170] Post-screen expansion of CleavER cells. Sorted GFP+ and GFP- KA2_CleavER_PM7 cells which were not processed for scRNA seq were expanded in RPM110 media for up to 2 weeks until a minimum of 2x106were available before cryopreservation and RNA extraction.
[0171] Antibody staining and panels: All antibodies were titrated on cells at a concentration of 1x106 / ml in cold separation buffer and incubated for 20mins at 4°C. Cells were washed once in cold separation buffer and resuspended at 1x106 / ml in cold separation buffer before FACS acquisition. The following panels were used (dilutions and cat numbers provided in brackets):
[0172] - TBNK+Nkp46: TBNK (1 : 10, CAT:644611),
[0173] - Memory phenotype staining: CD197-PE (1 :25, CAT:566741 ), CD45RA-BB515 (1 :100, CAT:564552), CD8-BV510 (1 :100, CAT:344732), CD4-APC (1 :50, CAT:565994), CD3- BV786 (1 :100, CAT:563800) CD56-BV650(1 :66, CAT:564057).
[0174] - CleaVER screens: CD8b-PE (1 :50, CATJM2217U), CD86-BV786 (1 :600, CAT:740990), CD3-BV786 (1 :200, CAT:563800), CD80-BV421 (1 :100, CAT:305221), CD137-PeCy7 (1 :100, CAT:309818), CD107a-APC (1 :100, CAT:641581), CD107a-APC (1 :33, CAT:560664), 7AAD (1 :40, CAT:420404).
[0175] ScRNA seq analysis. A single cell RNA-sequencing (scRNA-seq) analysis workflow has been written to recover tiles from scRNA-seq experiments performed on the BD Rhapsody. The workflow filters low quality read pairs based on the average phred score of bases and read length using cutadapt (Martin, 201 1). In the BD Rhapsody workflow, read 1 s hold cell barcode (cell label) and UMI (molecule label) information for each read pair. Read 1 s are filtered if they have an average base phred score of <20 or the read length is below 10bp. Read 2s hold the target sequences. These are also filtered based on an average base phred score of <20 or if the read length is below 50. The tile primers are then used as anchoring points to trim reads around the tile sequence. The remaining read pairs contain high confidence read 1 and 2 sequences. A fastp (Chen et al. 2018) report is then constructed to confirm the quality of the remaining read pairs. Following the read QC step, the cell barcodes in the read 1 s are extracted using the UMI-tools (Smith et al. 2017) package. A regular expression is used to capture the sequences of the cell barcodes and UMIs. These are then stored in the sequence identifier (first) line of the FASTQ. Putative cell barcodes are called by identifying barcodes within a hamming distance of 1 from each other and collapsing them into a single cell identifier. This is done to account for sequencing errors in the barcode which cause reads from one cell to splinter into many low read count cells. Once the cell barcodes have been extracted for each read pair, the reads are ready to be aligned. The BWA-mem aligner (Li & Durbin, 2009) is used to map the extracted FASTQs to a reference. Secondary alignments are filtered out, along with alignments which have a mismatch score of greater than one. After the alignment, reads are assigned to tiles using the featureCounts algorithm (Liao et al. 2014). Once reads have been assigned to tiles, UMI count correction is performed using UMI-tools (Smith et al. 2014). This is the process of counting the number of UMIs which have been assigned to each tile per cell. A tilexgene (tile-by-cell matrix) is obtained and the number of cells which express two or more UMIs of a tile are counted. This produces a count of the number of cells which support each tile. Tiles which are supported by fewer than five cells are removed to avoid spurious enrichments / depletions. Estimates of the proportion of cells in the positive cell fraction based on the raw count of cells which support a tile are obtained for each tile, and used to create a normalised rank by creating a Z-score distribution. This Z-score distribution is then used to identify tiles which have the highest levels of enrichment from the assay.
[0176] Sequences used in these examples are listed in Table 0 below.
[0177]
[0178] Table 0. Sequences used in examples. In SEQ ID NO: 6, the first 3 amino acids (ATM) are the start codon and Kozak sequence, followed by the 29 amino acids CMV peptide (starting “PP” and finishing “EF”), then the linker (starting “GS” and finishing “GP”, followed by the CFP sequence (239 amino acids). #=SEQ ID NO. Results
[0179] Although their use is not limited to this, the reporters of the present disclosure were developed in the context of the development of a high-throughput trimolecular TCR:peptide:MHC platform to identify immunogenic mutations. To screen for immunogenic neoantigens (iNeoAg), single cell RNA sequencing (scRNA-seq) will be performed on Physically Interacting Cells (PIC), a to identify an immunogenic triplet TCR:HLA:Neoantigen accompanied by a T-cell activation signature. PIC- SEQ (originally described in Giladi et al. 2020) is formed of two experimental procedures, a FACS- Based assay to identify PIC and scRNA-seq to reveal the interacting TCR and neoantigen. The reporters described herein were developed for use as a specific and sensitive T-cell activation signature. Identification of TCR:HLA:NeoAg triplet is dependent on the antigen expressing target being mono- allelic for a given HLA molecule, expressing a neoantigen and remaining intact after T-cell recognition to permit scRNA-seq. To optimise these experiments a healthy donor derived CMV- specific T-cell line recognising HLA-A*02:01 restricted epitope NLVPMVATV(NLV) was used as effector cells and the HLA-A*02:01 expressing cell line K562 (KA2) as antigen expressing targets. To confirm KA2 could elicit an Ag-specific T-cell response, KA2 were peptide pulsed or transfected with a tandem mini gene (TMG) encoding CMV 29mer amino acid sequence to measure endogenous processing and presentation of NLV T-cell epitope. TMGs were linked to CFP via a 2A sequence to measure TMG expression in KA2 (10% TMG expression). Target cells were then co-cultured with CMV-specific T-cells for 6hrs and immune response was measured according to degranulation marker CD107a and production of inflammatory cytokines IFN-y and TNF-a.
[0180] The high diversity of the mutated library that was designed for screening in this work enables the unique possibility to conduct an unbiased screen of 600,000 neoantigens for immunogenicity. Although already split into 6 pools of 100,000 TILE sequences, the low sensitivity of FACS based assays requires the enrichment of each library for immunogenic hits to permit detection in the triplet identifier assay (TIA). In this context, the inventors designed a novel antigen reporter cell line which utilises a degradation tag-13 (dTAG-13) system targeting mFKBP12, previously used for targetspecific protein degradation. The dTAG reporter construct used in the present examples contains a Green Fluorescent protein (GFP) linked with the dTAG-13 target mutant FKBP12 site (mFKBP12) via a Granzyme-B cleavage sequence (Figure 11). Treatment with dTAG-13, binds mFKBP12 and results in protein degradation of the tagged protein sequence. Here, the inventors expressed a dTAG based reporter as described herein in KA2 cells (referred to as KA2-dTAG or CleavER cells) and sorted for high expression of GFP. When antigen is expressed in CleavER cells and co-cultured with CD8T-cells, Ag-specific T-cells degranulate allowing Granzyme-B to directly enter the target cell through perforin pores. Following enzymatic cleavage by Granzyme- B, GFP will be released from mFKBP12 and upon dTAG-13 treatment will not be targeted for protein degradation remaining GFP positive (Figure 2A). CleavER cells not targeted by Ag- specific T-cells will not release GFP from mFKBP12 and therefore are negative for GFP expression after dTAG-13 treatment (Figure 2A).
[0181] Immunogenic neoantigen screening is a novel use of the dTAG system and so the inventors first determined the concentration of dTAG-13 and length of treatment required to acquire a GFP negative K562 cell line. CleavER cells were treated with increasing concentrations of dTAG-13 for 30mins, 1 hr, 2hrs and 4hrs at 37°C and subsequently assessed by FACS for GFP expression. Untreated CleavER cells were highly positive for GFP as expected and upon treatment by dTAG- 13, GFP expression was lost (Figure 6). The percentage of GFP loss was dependent on concentration of dTAG-13 and incubation time and 1 pM dTAG-13 with a 4hr incubation resulting in complete GFP loss. The inventors further observed that the cells stayed GFP negative over several days in the presence of dTAG-13 (data not shown).
[0182] Next, a proof-of-concept experiment was performed to determine if the reporters would signal antigen-specific recognition by CMV-specific T-cells. CleavER cells (K562 cells stably transduced with a GFP-GBC-FKBP12F36V construct where the GFP is driven from the CMV promoter in the PLX302 lentiviral backbone, where GBC= granzyme B cleavage sequence RIEAD; see Fig. 11) were peptide pulsed with CM V-de rived HLA-A*02:01 restricted epitope NLVPMVATV(NLV) or unpulsed, and incubated with CMV-specific T-cells for 4 hours or overnight (16hrs) at a 1 : 1 or 10:1 effectorstimulator ratio before subsequent treatment with 2uM dTAG-13 for 3hrs (Figure 7A). Cells were then measured for GFP expression via FACS as well as T-cell degranulation using CD107a expression. At both timepoints a high frequency of T-cells degranulated in response to peptide pulsed CleavER cells suggesting perforin and Granzyme-B was released (Figure 8). As before, CleavER cells alone demonstrated high expression of GFP which was lost upon treatment with dTAG-13, however incubation with 1 :1 ratio of CMV-specific T-cells with peptide pulsed targets rescued GFP loss at both 4hrs and 16hrs (Figures 8A, 8B). The 10:1 ratio of CMV-specific T-cells resulted in peptide-pulsed cells being killed entirely and so no GFP rescue was seen (Figure 8A, 8B). These data demonstrate antigen-specific, granzyme B mediated release of GFP from the mFKBP12 permitting detection of an antigen-reporter signal after engagement with antigen- specific T-cells. The cells pulsed with the peptides stayed positive over several days, and the data show a clear correspondence between the degranulation markers and the reporter signal. Figure 9 shows more detailed results for the 4 hrs co-culture with donor T-cells from a CMV+ donor at 1 :1 effector to sensor E:S ratio, followed by treatment with 1 uM dTAG-13 for 3hrs (same experiments as results in Figures 8A). Flow cytometry plots for GFP vs side scatter (SS) in the various conditions are shown, as well as summarised flow cytometry results.
[0183] Following this, the antigenic sensitivity of the reporters was determined. CleavER cells were either peptide pulsed with NLV peptide or transfected with a tandem mini gene (TMG) encoding CMV 29mer amino acid sequence to measure endogenous processing and presentation of NLV T-cell epitope. (TMG_NLV). TMG_NLV was linked to CFP via a 2A sequence to measure TMG expression in KA2. TMG_NLV expressing cells were then spiked into KA2_dTAG reporters at 50%, 20% and 5% and co-cultured with T-cells spiked with CMV-specific T-cells at similar frequencies, for 6hrs or 16hrs (Figure 8C). Cells were then treated with dTAG-13 for the final 3hrs of the coculture before subsequent FACS analysis for GFP, CD107a and activation marker CD137 expression. CleavER cells demonstrated increased GFP expression when compared with antigen negative cells when antigen was expressed and CMV-specific T-cells were present (Figure 8C). Furthermore, GFP expression was dependant on T-cell frequency and antigen frequency validating the antigen-specificity of the reporter (Figure 8C). However, although the frequency of degranulating T-cells increased during a 16hr incubation the frequency of GFP+ reporter cells did not change (Figure 8C, bottom). This is likely due to Ag-specific lysis of CleavER cells following CMV-specific T-cell recognition. Crucially, GFP expressing reporter cells were also positive for TMG_NLV expression showing dependence on antigen expression (Figure 8D). Furthermore, when T-cell and target cell doublets were gated GFP expressing cells showed increased CD137 expression demonstrating ag-specific doublet formation (Figure 8D). This confirms the CleavER cells are antigen-specific and therefore the reporter signal acts as a reliable signal for antigen enrichment.
[0184] Finally, the inventors determined if CleavER cells would report antigen recognition of or K562 TILE libraries. CleavER cells were transduced with a single TILE encoding the CMV epitope (CMV_TILE). CMV_TILE was then spiked into CleavER cells at decreasing frequencies of 100%, 50%, 20% and 5%and incubated with CMV-specific T-cells as described previously. Ag-specific T- cells were identified by CD107a or CD137 expression and at 6hrs incubation ag-specific T-cell responses were seen against CMV_TILE expressing cells (Figure 10A, B) albeit at low frequencies despite 100% expression of CMV-TILE. At 16hrs incubation Ag-specific T-cell responses were noticeably increased against CMV_TILE suggesting TILE expression is lower compared with TMG_NLV (Figure 8C, D) and longer incubation times are advantageous to detect a lower frequency T-cell response against TILE libraries. At 16hrs, reporter activity was observed, above background levels, when as low as 5% CMV_TILE bearing cells were present (Figure 10). The data on Figure 10 show that the assay has high sensitivity, and is able to detect antigen presenting cells representing as low as 5% of the target cell population using a T cell population in which more than 5% of T cells recognise the antigen, and is also able to identify reactive T cells representing as low as 5% of the T cell population using a target cell population in which the cognate antigen presenting cells represent more than 5% of the target cell population.
[0185] Work is currently underway to generate a brighter, more stable reporter using mStayGOLD fluorescent protein. This is expected to further improve the sensitivity of the assay. Furthermore, improving sensitivity by increasing TILE expression using alternative promoters is also being done simultaneously. Nevertheless, these data demonstrated CleavER cells reporters using GFP can be used to detect TILE library responses. GFP positive CleavER cells can be FACs sorted and cultured to expand cell numbers to align with the limit of detection of the triplet identifier assay using PIC-SEQ.
[0186] To ultimately demonstrate CMV_TILE enrichment from CleavER cells the inventors have performed an experiment whereby CleavER cells expressing CMV_TILE was spiked into CleavER cells at 20% and co-cultured with CMV-specific T-cells. After 16hrs, GFP positive reporter cells were sorted by FACS and run through the BDRhapsody™ scRNA sequencing platform to assess if CMV_TILE has been enriched from a diverse TILE Library (Figure 7B). In particular, 20 million T-cells and 20M million CleavER cells were co-cultured for 16hr (followed by a 3hr DTAG treatment). Cells were then sorted between a CD3- CD86+ GFP- fraction and a CD3- CD86+ GFP+ fraction (gating on CD86+ selects the K562 cells, T cells express CD3 so sorting for CD3- excludes T cells, and GFP- / GFP+ fractions separates respectively the negative and positive cells). The former yielded 100k cells for scRNA sequencing, and the latter yielded 57848 cells that were processed directly for scRNA sequencing or 140k cells after sorting and expansion in culture. The cells were subject to scRNAseq using the BD Rhapsody platform (GFP+:22557 cells, GFP-: 23466 cells), as well as cDNA synthesis, PCR to expand the tiles and sequencing. Fig. 8E shows results of these experiments. scRNA analysis (unique molecular identifiers (UMI) counts for the CMV tile) of GFP positive and GFP negative CleavER cells spiked with 20% CMV antigen expressing CleavER cells following co-culture with CMV-specific T-cells are shown. The data show a Six fold increase in the % of cells with significant CMV expression in the GFP positive fraction compared to the GFP negative fraction, demonstrating a strong CMV antigen enrichment on GFP+ CleavER cells. Note that the data shown is for a proof-of-concept experiment with limited read depth, such that even better results would be expected with increased read depth. This work demonstrates that the reporter is usable as a reporter at the single cell level (in addition to the bulk data demonstrated on Figures 8A-D), thereby demonstrating that the reporter can be used for sensitive identification of antigens and / or TCRs in large scale single cell screens.
[0187] In summary, the preset examples demonstrate a versatile reporter system, in which each of the degron, selection marker and cleavage sequence can be adapted to suit a particular assay. The present examples demonstrate the use of this reporter in the context of identification of T cell activation events and associated antigens and / or TCRs, in a sensitive manner that is amenable to high throughput screening.
[0188] Example 2
[0189] Figure 12 demonstrates CleavERs ability to detect immunogenic antigenic hits as low as 2% expression frequency following bulk RNA-seq analysis of GFP+ CleavER cells. Here, 5x106polyclonally expanded T-cells were spiked with 2% CMV-specific T-cells and co-cultured with KA2_CLEAVER_PM7 targets (20K neoantigen library) spiked with CMVTILE expressing KA2_CLEAVER (KA2_CLEAVER_CMV) at 2% (E:T ratio 1 : 1). Cells were co-cultured for 16hrs at 37°C before treatment with DTAG-13 for 3hrs and subsequently stained for CD80 (CleavER targets), CD3 (T-cells) and CD137 (T-cell Activation marker). As a control for autofluorescence influence on hit calling post DTAG- 13 treatment, KA2_CleavER without T-cells were also prepared and CleavER cells were FACS sorted for CD80+CD3- GFPpos and GFPneg populations. Sorted cells were expanded for 7 days to reach 2x106 cells required for the bulk RNA seq pipeline and lysed using RLTIysis buffer. DNA libraries were prepared and sequenced via Nextseq 2000 and data were processed using a DECOD bioinformatic pipeline for hit calling. The No T-cell CleavER control was used to determine a threshold for Hit calling and CMVTILE was identified as significantly enriched at 2% antigen frequency (Figure 12).
[0190] Example 3
[0191] To demonstrate that CleavER would identify novel immunogenic hits from a diverse library, we screened a 20k Neoantigen Printed mutant library 7 (PM7) against healthy donor T-cells. 20x106PM7-expanded T-cells from 3 HLA-A*02:01 + healthy donors were co-cultured at a 1 :1 E:T ratio with KA2_CleavER_GFP_PM7 target cells overnight and subsequently treated with dTAG-13 for 3hrs. Cocultures were then FACS sorted for GFPpos and GFPneg target cell populations processed via scRNA BD Rhapsody workflow to identify TILE expression following an Empirical Bayes Estimation workflow. In brief, Candidate TILEs identified in a cell were considered hits if it had a minimum UMI support of 2 and the enrichment across cells was measured as a frequency of the GFPpos population compared to GFPneg populations. Enriched TILE expression in GFPpos cells was found in all donor screens against PM7 library (Figure 13). In total, 1 12 hits were identified from H013 CleavER screen, 103 hits were identified from H011 CleavER screen and 94 hits were identified in H018 CleavER screen (Figure 13). We demonstrate CleavER can identify novel immunogenic Neoantigen sequences from a diverse antigen library containing 20k Neoantigen candidate TILE sequences.
[0192] Example 4
[0193] Single cell cloning (SCC) of CleavER cells improves CleavER performance by removing high background reporter signal following DTAG-13 treatment, caused by heterogenous reporter expression in a bulk population. In brief, GFP-CleavER cassettes were lentivirally transduced into K562 cells expressing HLA-A*02:01 (KA2_C LEAVER). HLA-A2hi, CD80hi, CD86hi and GFPhi expressing KA2_CleavER were then single cell sorted into 96 well plates and expanded for 2 weeks. Expanded SCC were then screened by FACS for the largest fold change in GFP expression following a 4hr DTAG-13 treatment. In total, 52 SCC were screened and SCC B5 (KA2_CLEAVER_GFP_B5) was found to have the biggest fold difference in GFP expression following DTAG-13 treatment (Figure 15A). Next we compared the original bulk KA2_CleavER_GFP to the SCC_B5 and determined if prolonged DTAG-13 treatment time removed residual GFP signal post DTAG treatment. SCC_B5 demonstrated a much improved GFP negative signal post DTAG-13 compared to the bulk at short treatment times (3-4hrs). Increased DTAG-13 treatment times (24-48hrs) reduced GFP signal in the bulk, however, SCC_B5 consistently demonstrated a cleaner background GFP signal (Figure 15B&C). Furthermore GFP expression is homogenous and less diverse than the bulk cell line. This data shows the advantages of single cell cloning to achieve a reliable, consistent CleavER cell line suitable for low frequency immunogenic antigen detection.
[0194] Example 5
[0195] Reporter protein element.
[0196] Here, we demonstrated functionality of alternative modular elements in CleavER cells to demonstrate versatility of the antigenic reporter. First, we switched the GFP fluorescent reporter protein for 2 other fluorescent proteins, mCherry or Staygold in the cleavER construct. Next, we expressed them lentivirally in HLA-A*02:01 + K562 and generated SCC as before and screened for the largest fold difference in fluorescent signal post DTAG-13 treatment. In total, 6X KA2_CLEAVER_mcherry SCC and 1x KA2_CLEAVER_StayGold SCC (Figure 16) were selected based on this criteria. SCC were then tested in DTAG-13 kinetic assays to determine optimum treatment time and the cleanest negative signal for each reporter protein following DTAG-13 treatment. 24hr DTAG-13 treatment improved background signal more thoroughly in all SCC tested with the alternative reporter proteins.
[0197] Cell surface protein reporter element
[0198] Fluorescent proteins require FACS sorting of CleavER cells to identify the immunogenic antigen recognised by T-cells using Next generation sequencing. An alternative approach would be to use a reporter protein expressed on the cell surface that would permit magnetic enrichment. We similarly switched the GFP reporter protein with a chimeric protein containing anti-CD34 and anti- CD20 minimal epitopes (RQR8) or a modified NGFR protein containing a CD8 intracellular tail (NGFR8). Each construct was lentivirally introduced into K562 expressing HLA-A*02:01 and protein expression was assessed following DTAG-13 treatment, using fluorescently labelled antibodies against CD34 or CD271 (NGFR). We demonstrated both RQR8 and NGFR8 were degraded following DTAG-13 treatment but with differing kinetics. RQR8_CLEAVER demonstrated loss of CD34 expression when treated for 19hrs with DTAG-13 (Figure 17), and NGFR8_CLEAVER demonstrated a slower degradation time (Figure 18). Instead, the overall amount of NGFR protein (Geomean) was reduced to half the signal by 72hrs DTAG-13 treatment (Figure 18C). Some cells demonstrated complete loss of NGFR at 72hrs and so represent heterogeneity within a bulk CLEAVER population (Figure 18B). Single cell clones were generated of NGFR8_CLEAVER to identify clones with complete NGFR degradation and we identified 11 SCC (Figure 18D) which showed a cleaner NGFR degradation signal (Figure 18E).
[0199] Protein Degradation element.
[0200] Here, we switched the protein degradation module from DTAG-13 binding Mfkb12 to the fast acting auxin-inducible degron tag (mAID) in an alternative GFP CleavER construct. This is a two-protein component system, so two genetic modifications are required for degron activity. The protein of interest was fused with a 7-kD degron, called mini-AID (mAID) and OsTIRI (TIR1 derived from Oryza sativa) was expressed to form an E3 SKP1-CUL1-F-box ligase, SCF-OsTIR1 (also called CRL1-OsTIR1). The GFP_Auxin_CleavER construct was expressed in HLA-A*02:01 + K562 and treated with increasing doses of indole-3-acetic acid (IAA; a natural auxin) for increasing timepoints. We demonstrated the highest dose of IAA permitted complete GFP degradation after 4hrs IAA treatment comparable to DTAG-13 based cleavERs (see Figure 19).
[0201] In conclusion, the versatility of CleavER provides advantages. Each modular element i.e, reporter protein, enzymatic cleavage site or DEGRON sequence, can be easily switched out depending on the user’s needs.
[0202] Example 6
[0203] Here, we generated an HLA negative CleavER SCC GFP reporter which can be transduced with any HLA-allele of interest, permitting an easily customisable reporter to screen for antigen hits across multiple HLA types. This would permit an off-the-shelf screening platform which could be used for example: in patient HLA-allele matched screens of immunogenic antigen identification for vaccine design, or antigen-specific T-cell products. To achieve this, we transduced wildtype K562 obtained from public health England (ECACC 89121407) and lentivirally expressed CD80 and CD86 T-cell costimulatory proteins and knocked out LAMP1 (encoding CD107a) using CRISPR- Cas9 and SCC for CD80hiCD86hiCD107aneg. Next, GFP_GZB_Mfkb12 Cleaver cassette was transduced and GFPpositive cells were similarly SCC for CD80hiCD86hiCD107aneg and screened for the biggest fold difference in GFP signal following DTAG-13 treatment.
[0204] New antigenic libraries can be screened in different HLA types using the DNA sequence of the HLA-allele of interest (https: / / www.ebi. ac.uk / ipd / imgt / hla / alleles / allele / ?accession=HLA00005). This will be cloned into a lentiviral expression plasmid which can be used to generate virus and transduced into the ready to go HLA negative CleavER SCC. Mono-allelic HLA bearing CleavER cells are then enriched using magnetic isolation and further transduced with a specific antigen TILE library which can be purified using classical antibiotic selection techniques (such as puromycin).
[0205] Example 7
[0206] T-cells induce target cell specific killing following release of granzyme-B and perforin upon T-cell recognition of expressed antigen. Therefore, in positive selection screens such as CleavER, cells will have a finite reporting window before ultimately succumbing to cell death. Generating a Granzyme-B resistant cell line aims to maximise the application of CleavER by preserving CleavER cells recognised by T-cells. This would likely permit the identification of more immunogenic hits by next generation sequencing. Furthermore, CleavER cells recognised by T-cells could be sorted and expanded for downstream applications such as immunepeptidomics to identify the minimal epitope or used in T-cell expansion protocols to enrich for reactive T-cells.
[0207] References
[0208] Hirano, M., Ando, R., Shimozono, S. et al. A highly photostable and bright green fluorescent protein. Nat Biotechnol 40, 1 132-1 142 (2022). Liesche C, et al. Single-Fluorescent Protein Reporters Allow Parallel Quantification of Natural Killer Cell-Mediated Granzyme and Caspase Activities in Single Target Cells. Front Immunol. 2018 Aug 8;9:1840.
[0209] Kula T, et al. T-Scan: A Genome-wide Method for the Systematic Discovery of T Cell Epitopes. Cell. 2019 Aug 8;178(4):1016-1028.e13.
[0210] Miyamoto DK, Curnutt NM, Park SM, Stavropoulos A, Kharas MG, Woo CM. Design and Development of IKZF2 and CK1 a Dual Degraders. J Med Chem. 2023 Dec 28;66(24): 16953- 16979.
[0211] Yesbolatova, A., Saito, Y., Kitamoto, N. et al. The auxin-inducible degron 2 technology provides sharp degradation control in yeast, mammalian cells, and mice. Nat Commun 1 1 , 5701 (2020).
[0212] Nishimura, K., et al. An auxin-based degron system for the rapid depletion of proteins in nonplant cells. Nat. Methods 6, 917-922 (2009).
[0213] Nabet B, et al. The dTAG system for immediate and target-specific protein degradation. Nat Chem Biol. 2018 May;14(5):431 -441.
[0214] Shukla SA, et al. Comprehensive analysis of cancer-associated somatic mutations in class I HLA genes. Nat Biotechnol. 2015 Nov;33(11):1 152-8.
[0215] Andras Szolek, et al. OptiType: precision HLA typing from next-generation sequencing data, Bioinformatics, Volume 30, Issue 23, December 2014, Pages 3310-3316
[0216] Giladi A, et al. Dissecting cellular crosstalk by sequencing physically interacting cells. Nat Biotechnol. 2020 May;38(5):629-637.
[0217] O'Donnell TJ, Rubinsteyn A, Laserson U. MHCflurry 2.0: Improved Pan-Allele Prediction of MHC Class l-Presented Peptides by Incorporating Antigen Processing. Cell Syst. 2020 Jul 22;1 1 (1 ):42- 48. e7.
[0218] B Fehse, A Richters, K Putimtseva-Scharf, et al. CD34 splice variant: an attractive marker for selection of gene-modified cells. Mol. Ther, 1 (5 Pt 1) (2000), pp. 448-456.
[0219] Philip B, et al. A highly compact epitope-based marker / suicide gene for easier and safer T-cell therapy. Blood. 2014 Aug 21 ; 124(8): 1277-87.
[0220] A Di Stasi, SK Tey, G Dotti, et al. Inducible apoptosis as a safety switch for adoptive cell therapy. N Engl J Med, 365 (18) (201 1 ), pp. 1673-1683.
[0221] Rudoll T, Phillips K, Lee SW, Hull S, Gaspar O, Sucgang N, Gilboa E, Smith C. High-efficiency retroviral vector mediated gene transfer into human peripheral blood CD4+ T lymphocytes. Gene Ther. 1996 Aug;3(8):695-705.
[0222] Jan M, et al. Reversible ON- and OFF-switch chimeric antigen receptors controlled by lenalidomide. Sci Transl Med. 2021 Jan 6;13(575):eabb6295.
[0223] Chung HK, Jacobs CL, Huo Y, Yang J, Krumm SA, Plemper RK, Tsien RY, Lin MZ. Tunable and reversible drug control of protein production via a self-excising degron. Nat Chem Biol. 2015 Sep; 11(9)713-20.
[0224] Christina Backes, et al., GraBCas: a bioinformatics tool for score-based prediction of Caspase- and Granzyme B-cleavage sites in protein sequences, Nucleic Acids Research, Volume 33, Issue suppl_2, 1 July 2005, Pages W208-W213. Li, X., Song, Y. Proteolysis-targeting chimera (PROTAC) for targeted protein degradation and cancer therapy. J Hematol Oncol 13, 50 (2020).
[0225] Yan, Y., Xu, TH., Melcher, K. et al. Defining the minimum substrate and charge recognition model of gamma-secretase. Acta Pharmacol Sin 38, 1412-1424 (2017).
[0226] MARTIN, Marcel. Cutadapt removes adapter sequences from high-throughput sequencing reads. EMBnet.journal, [S.I.], v. 17, n. 1 , p. pp. 10-12, may 2011. ISSN 2226-6089.
[0227] Yang Liao, Gordon K. Smyth, Wei Shi, featurecounts: an efficient general purpose program for assigning sequence reads to genomic features, Bioinformatics, Volume 30, Issue 7, April 2014, Pages 923-930.
[0228] Heng Li, Richard Durbin, Fast and accurate short read alignment with Burrows-Wheeler transform, Bioinformatics, Volume 25, Issue 14, July 2009, Pages 1754-1760.
[0229] Smith T, Heger A, Sudbery I. UMI-tools: modeling sequencing errors in Unique Molecular Identifiers to improve quantification accuracy. Genome Res. 2017 Mar;27(3):491-499.
[0230] Shifu Chen, Yanqing Zhou, Yaru Chen, Jia Gu, fastp: an ultra-fast all-in-one FASTQ preprocessor, Bioinformatics, Volume 34, Issue 17, September 2018, Pages i884— i890.
[0231] All references cited herein are incorporated herein by reference in their entirety and for all purposes to the same extent as if each individual publication or patent or patent application was specifically and individually indicated to be incorporated by reference in its entirety.
[0232] The specific embodiments described herein are offered by way of example, not byway of limitation. Various modifications and variations of the described compositions, methods, and uses of the technology will be apparent to those skilled in the art without departing from the scope and spirit of the technology as described. Any sub-titles herein are included for convenience only and are not to be construed as limiting the disclosure in any way. Unless context dictates otherwise, the descriptions and definitions of the features set out above are not limited to any particular aspect or embodiment of the invention and apply equally to all aspects and embodiments which are described. Throughout the specification and claims, the following terms take the meanings explicitly associated herein, unless the context clearly dictates otherwise. The phrase “in one embodiment” as used herein does not necessarily refer to the same embodiment, though it may. Furthermore, the phrase “in another embodiment” as used herein does not necessarily refer to a different embodiment, although it may. Thus, as described below, various embodiments of the invention may be readily combined, without departing from the scope or spirit of the invention. It must be noted that, as used in the specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from “about” one particular value, and / or to “about” another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by the use of the antecedent “about,” it will be understood that the particular value forms another embodiment. The term “about” in relation to a numerical value is optional and means for example + / - 10%. Throughout this specification, including the claims which follow, unless the context requires otherwise, the word “comprise” and “include”, and variations such as “comprises”, “comprising”, and “including” will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps. Other aspects and embodiments of the invention provide the aspects and embodiments described above with the term “comprising” replaced by the term “consisting of’ or “consisting essentially of’, unless the context dictates otherwise, “and / or” where used herein is to be taken as specific disclosure of each of the two specified features or components with or without the other. For example “A and / or B” is to be taken as specific disclosure of each of (i) A, (ii) B and (iii) A and B, just as if each is set out individually herein.
[0233] The features disclosed in the foregoing description, or in the following claims, or in the accompanying drawings, expressed in their specific forms or in terms of a means for performing the disclosed function, or a method or process for obtaining the disclosed results, as appropriate, may, separately, or in any combination of such features, be utilised for realising the invention in diverse forms thereof.
Claims
CLAIMS1. An isolated nucleic acid encoding a reporter fusion protein comprising a selection marker, a protease recognition sequence and a degron.
2. The isolated nucleic acid of claim 1 , comprising a sequence encoding a selection marker in frame with a sequence encoding a protease recognition sequence and a sequence encoding a degron.
3. The isolated nucleic acid of claim 1 or claim 2, wherein the degron is a protein or peptide that leads to inducible degradation of the degron containing fusion protein; and / or wherein the degron is a peptide or protein that targets the reporter fusion protein for proteasomal degradation in the presence of a degradation ligand, optionally wherein the degron is a peptide or protein that binds to a degradation ligand, and the degradation ligand is a molecule that binds to the degron and a ubiquitin ligase.
4. The isolated nucleic acid of claim 3, wherein the degron is a mFKB12 peptide and the degradation ligand is a dTAG molecule, optionally dTAG-13; wherein the degron is an IKZF2 peptide, an IKZF3 degron tag, or a CK1 a protein or peptide and the degradation ligand is lenalidomide or a lenalidomide analogue; wherein the degron is an auxin-inducible degron and the degradation ligand is an auxin or auxin analogue; wherein the degron is a Small Molecule-Assisted Shutoff (SMASh) tag and the degradation ligand is an inhibitorof a protease included in the SMASh tag, optionally an inhibitor of NS3 protease; wherein the degron is a degron of a SMASh tag, optionally a NS4A peptide; wherein the degron is a proteolysis targeting chimera (PROTAC) target protein or peptide and the degradation ligand is a PROTAC comprising a ligand for the target protein or peptide; or wherein the degron is a HaloTag and the degradation ligand is a HaloPROTAC.
5. The isolated nucleic acid of any preceding claim, wherein the selection marker is a positive selection marker, wherein the selection marker comprises a fluorescent marker, wherein the selection marker comprises a cell surface marker, and / or wherein the selection marker comprises a cell survival protein.
6. The isolated nucleic acid of claim 5, wherein the fluorescent marker is a fluorescent protein, optionally selected from GFP, mCherry and mStayGold, and / orwherein the cell surface marker is a protein that is expressed on the surface of cells, optionally a CD34 protein, a CD19 protein, a NGFR protein, or a PD-L1 protein and / or wherein the cell survival protein is an antiapoptotic protein, optionally selected from BCL2, MCL1 and XIAP, an immune suppression protein,optionally PD-L1 , or a mutant protein that confers resistance to a cytotoxic compound or composition, optionally a mutant BCR-ABL protein.
7. The isolated nucleic acid of any preceding claim, wherein the protease recognition sequence is an amino acid sequence that is specifically recognised by a protease selected from granzyme B, gamma secretase, and a caspase, optionally caspase 3, 7, 8 and / or 10, optionally wherein:(a) the protease is granzyme B and the protease cleavage sequence is a sequence comprising amino acid sequence IEXD (SEQ ID NO: 23) where X can be any amino acid, VEXD (SEQ ID NO: 24) where X can be any amino acid, IGXD (SEQ ID NO: 25) where X can be any amino acid, or VGXD (SEQ ID NO: 26) where X can be any amino acid,(b) the protease is caspase-3 / 7 and the protease cleavage sequence is a sequence comprising amino acid sequence DXXD (SEQ ID NO: 18) or DEXD (SEQ ID NO: 27) where X can be any amino acid,(c) the protease is caspase-8 / 10 and the protease cleavage sequence is a sequence comprising amino acid sequence LQTD (SEQ ID NO: 4), LETD (SEQ ID NO: 29), VETD (SEQ ID NO: 30), LEHD (SEQ ID NO: 31), or VEHD (SEQ ID NO: 32), or(d) the protease is gamma secretase and the protease cleavage sequence is a sequence comprising amino acid sequence EDVGSNKGAIIGLMVGGWIATVIVITLVMLKKK (SEQ ID NO: 33), FMYVAAAAFVLLFFVGCGVLL (SEQ ID NO: 34), LLYLLAVAWIILFIILLGVI (SEQ ID NO: 35), LPLLVAGAVLLLVILVLGVMV (SEQ ID NO: 36) or PVLCSPVAGVILLALGALLVL (SEQ ID NO: 37), or a homologue or variant thereof that is recognised by gamma secretase.
8. The isolated nucleic acid of any preceding claim, wherein the protease recognition sequence is a peptide comprising a protease recognition motif, wherein the protease recognition motif is exposed in the reporter fusion protein, optionally wherein the protease recognition sequence forms an exposed loop comprising the sequence recognition motif.
9. The isolated nucleic acid of any preceding claim, wherein the protease recognition sequence comprises a portion of the sequence of a native substrate of the protease, said portion comprising the protease recognition motif of the protease, or a sequence having at least 80%, 85%, 90%, or 95% sequence identity to said sequence provided that the protease recognition motif is maintained, wherein the protease is granzyme B and the native substrate is a BID protein, or wherein the protease is caspase-3 / 7 and the native substrate is a PARP1 protein, and / or wherein the protease recognition sequence comprises a portion of the sequence of human BID comprising amino acids Leu47 to Ser78, or a homologue thereof or sequence having at least 80%, 85%, 90%, or 95% sequence identity to said sequence, or wherein the protease recognition sequence comprises a portion of the sequence of human PARP1 comprising amino acids Val202 and Asp217, or ahomologue thereof or sequence having at least 80%, 85%, 90%, or 95% sequence identity to said sequence.
10. A vector comprising the isolated nucleic acid of any preceding claim, optionally wherein:(a) the vector is a viral vector, optionally a lentiviral vector; and / or(b) the vector, further comprises one or more promoter sequences, optionally selected from: a CMV promoter, SFFV promoter and a hPGK promoter; and / or(c)the vector further comprises one or more antibiotic resistance genes and / or one or more candidate antigen sequences.11 . A reporter fusion protein encoded by the isolated nucleic acid of any of claims 1 to 9 or expressed from the vector of claim 10.
12. A cell comprising the vector of claim 10, the nucleic acid of any of claims 1 to 9, or the reporter fusion protein of claim 1 1 , optionally wherein the cell is an antigen presenting cell and / or a cancer cell, and / or wherein the cell is a MHC monoallelic cell or an HLA negative cell, and / or wherein the cell is an immortalised cell, optionally wherein the immortalised cell is from a cancer cell line, optionally a leukaemia cell line, optionally a k562 cell line and / or is from an HLA negative cell line.
13. A library of cells according to claim 12, wherein the cells in the library together display a plurality of antigens from a library of antigens, optionally wherein the plurality of antigens comprise at least 10, at least 20, at least 50, at least 100, at least 1000, at least 5000 or at least 10000 different antigen peptides and / orwherein the cells in the library are cells that have been modified to express one or more antigens of said plurality of antigens.
14. A method of making a reporter cell or reporter cell library, the method comprising obtaining a cell or population of cells and genetically modifying said cell(s) to express a reporter fusion protein according to claim 1 1 , optionally wherein the method further comprises the step of single cell cloning and / orwherein the reporter cell or library has the features of any of claims 12 or 13.
15. A kit comprising: (i) an isolated nucleic acid according to claims 1 to 9, a vector according to claim 10, or a cell or cell library according to claims 12 or 13; and (ii) a degradation ligand associated with the reporter fusion protein, wherein a degradation ligand associated with the reporter fusion protein is a compound that binds to the degron and causes the degradation of the reporter fusion protein.
16. A kit comprising: (i) an isolated nucleic acid according to claims 1 to 9 or a vector according to claim 10; and (ii) one or more cells, optionally wherein the one or more cells have the features of any of claims 13 or 14.
17. The kit of claim 16, further comprising: (ii) a degradation ligand associated with the reporter fusion protein, wherein a degradation ligand associated with the reporter fusion protein is a compound that binds to the degron and causes the degradation of the reporter fusion protein.
18. The kit of claim 15 or claim 17, wherein the degradation ligand is a molecule that binds to the degron and a ubiquitin ligase; optionally wherein the degron is a mFKB12 peptide and the degradation ligand is a dTAG molecule, optionally dTAG-13; wherein the degron is the degron is an IKZF2 peptide, an IKFZ3 degron tag, or a CK1 a protein or peptide and the degradation ligand is lenalidomide or a lenalidomide analogue; or wherein the degron is an auxin-inducible degron and the degradation ligand is an auxin or auxin analogue; wherein the degron is a Small Molecule- Assisted Shutoff (SMASh) tag and the degradation ligand is an inhibitor of a protease included in the SMASh tag, optionally an inhibitor of NS3 protease; wherein the degron is a degron of a SMASh tag, optionally a NS4A peptide; wherein the degron is a proteolysis targeting chimera (PROTAC) target protein or peptide and the degradation ligand is a PROTAC comprising a ligand for the target protein or peptide; or wherein the degron is a HaloTag and the degradation ligand is a HaloPROTAC.
19. A method of isolating an immunogenic peptide and / or a T cell receptor (TCR) reactive to a peptide, the method comprising: incubating a population of cells or a library of cells according to any of claims 12 or13 in the presence of a population of T cells and optionally a degradation ligand associated with the reporter fusion protein, the population of cells comprising cells presenting one or more of a plurality of candidate peptides on their cell surface in complex with a MHC molecule, optionally wherein the MHC molecule is encoded by a selected HLA allele; selecting one or more doublets of cells that comprise a T cell and an antigen presenting cell using a selection method associated with the selection marker comprised in the reporter fusion protein, thereby isolating an immunogenic peptide and a T cell expressing a TCR reactive to the peptide in the context of the selected HLA allele.
20. The method of claim 19, further comprising sequencing nucleic acid material isolated from one or more of the selected doublets, thereby identifying a candidate peptide associated with the antigen presenting cell in the doublet as an immunogenic peptide and / or a TCR expressed by the T cell in the doublet as a reactive TCR, optionally wherein the sequencing is single cell RNA sequencing.21 . The method of claim 18 or claim 19, wherein the selecting one or more doublets of cells that comprise a T cell and an antigen presenting cell is performed using a cell sorting assay, optionally fluorescence activated cell sorting (FACS) or magnetic cell sorting (MACS), and / or wherein the incubating is performed for at most 24 hours, between 12 and 24 hours, about 16 hours, between 4 and 12 hours, or about 4 hours, optionally comprising or followed by incubation for at least 2 hours or at least 3 hours in the presence of the degradation ligand.
22. The method of any of claims 18 to 21 , wherein the cells or library of cells have been peptide pulsed with the one or more of the plurality of candidate peptides, or wherein the cells or library of cells have been modified to express the one or more of the plurality of candidate peptides, optionally wherein the method comprises introducing an expression vector in cells of the population of antigen presenting cells prior to the incubating, the expression vector comprising a sequence encoding for the peptide, optionally wherein the expression vector is a lentiviral expression vector; and / or wherein the population of antigen presenting cells have been modified to include one or more expression vectors from a library of expression vectors, each expression vector of the library comprising a sequence encoding for a respective peptide and optionally a barcode sequence, wherein the population of immortalised cells comprise cells that express at least 10, at least 20, at least 50, at least 100, at least 1000, at least 5000 or at least 10000 different peptides.
23. A method of screening for a set of conditions to identify conditions that affect antigen recognition, the method comprising:Performing the method of any of claims 18 to 22using a reference condition and a condition of the set of conditions to be screened; andComparing the number and / or identity of the selected doublets, wherein differences in the number and / or identity of the selected doublets between the reference condition and the test condition is indicative of the test condition affecting antigen recognition, optionally wherein the condition is an incubation condition and / or a condition that the antigen presenting cells and / or T cells have been subjected to prior to incubation, optionally wherein the condition is exposure to an active compound or composition.
24. A method of training an immunogenicity prediction algorithm, wherein an immunogenicity prediction algorithm comprises a machine learning model that takes as input an antigen sequence and optionally a TCR sequence and / or MHC molecule sequence or pseudosequence and produces as output an indication of whether the antigen is likely to be immunogenic, the method comprising:screening a library of candidate antigens for antigens that are immunogenic using a method according to any of claims 18 to 23, thereby obtaining training data comprising candidate antigens of the library of candidate antigens and associated labels indicating whether each candidate antigen was identified as immunogenic, immunogenic triplets each comprising a candidate peptide of a library of candidate peptides, an MHC molecule and a TCR, and / or pairs of immunogenic antigens and associated TCRs; and training the machine learning model using said training data.
25. A method of determining the effect of one or more test conditions on apoptosis of a target cell population or gamma secretase activity in a target population, the method comprising: obtaining a target cell population comprising cells as claimed in claim 12, wherein the protease cleavage sequence is a caspase recognition sequence; culturing the target cells in the presence off the one or more test conditions, and optionally the degradation ligand associated with the reporter fusion protein; selecting or selectively quantifying cells of the target cell population that express the reporter fusion protein using a selection method associated with the selection marker comprised in the reporter fusion protein; and comparing the selected cells or quantities of selected cells between the one or more test conditions and / or between the one or more test conditions and corresponding reference values, thereby determining the effect of the respective test conditions on apoptosis of the target cells, optionally wherein the one or more test conditions are selected from: exposure to one or more compounds or compositions, co-culture with one or more cell populations, and exposure to one or more physico-chemical perturbations.
Citation Information
Patent Citations
Improved adapter for mechanically coupling a pump and a prime mover
EP0964998A1
Method for treating cancer
WO2016174085A1
Tunable endogenous protein degradation
WO2017024319A1
Methods and compositions for identifying epitopes
WO2018227091A1
Identification of clonal neoantigens and uses thereof
WO2022207925A1