Method for identifying temporal distinct cell lineages following a cellular event
The integration of CCA with time-linked data and protein expression profiles addresses the limitations of existing methods by accurately tracking cell lineage dynamics and differentiation, offering detailed insights into cellular activation and development.
Patent Information
- Application Number
- PCT/EP2025/064354
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-24
- Filing Date
- 2025-05-23
- Publication Date
- 2025-11-27
AI Technical Summary
Current methods for identifying and analyzing cell lineages following cellular events, such as cell signaling, fail to capture the dynamic nature of gene and protein expression changes over time, leading to ambiguities in differentiating true signals and autofluorescence, which hampers high-dimensional analysis.
A novel method combining Canonical Correspondence Analysis (CCA) with time-linked expression data and multidimensional protein profiles to classify and group cells based on their activation and differentiation status, using fluorescent timer proteins to link protein expression to time elapsed since the cellular event.
Enables precise identification of temporally distinct cell lineages and their functional status by analyzing protein expression changes over time, providing comprehensive insights into cellular trajectories and differentiation processes.
Smart Images

Figure EP2025064354_27112025_PF_FP_ABST
Abstract
Description
[0001] METHOD FOR IDENTIFYING TEMPORAL DISTINCT CELL LINEAGES FOLLOWING A CELLULAR EVENT
[0002] Field of the invention
[0003] The present invention relates to a computer-implemented method for identifying temporally distinct cell lineages following a cellular event such as a cell signalling event, as well as systems and computer-readable storage mediums for implementing the method.
[0004] Background to the invention
[0005] The fields of cellular biology, which include developmental cell biology, stem cell biology, immunology, and neuroscience, encompass the study of cellular events and subsequent cell lineage development and differentiation. The investigation of temporally dynamic changes of cells under development or differentiation is key for improving the understanding of cellular processes and events. Traditional approaches for identifying and grouping cells based on activation and differentiation status typically use methods including flow cytometry, microscopy, and sequencing analysis. However, without the use of an advanced technology, experimental measurements produce no more than static snapshots of the expression profiles of proteins or genes, which fail to capture the dynamic nature of cellular responses following events such as cell signalling. There exists a critical need for advanced technological approaches that can elucidate the gradual changes in gene (protein) expression as cells mature into various lineages, pinpoint key cellular stages pivotal to specific cellular events, and unravel the specific genes or proteins implicated at these stages (Holguera, I. & Desplan, C., 2018; Raimundo, J. & Levine, M., 2023; Voortman, L. et al., 2022).
[0006] The development and differentiation of cells are understood through the expression of a set of proteins (genes), which characterize each developmental / differentiation step. Recent advancements in flow cytometry and high-throughput sequencing technologies have enabled the simultaneous analysis of many genes or proteins, thereby providing a sophisticated understanding of the developmental / differentiation stages of cells in a detailed manner. For this purpose, multidimensional analysis is essential, and methods such as Principal Component Analysis (PCA), Uniform Manifold Approximation and Projection (UMAP), or t-distributed Stochastic Neighbour Embedding (t-SNE) are often employed. The combination of PCA and clustering algorithms such as k-means has been used to identify cell populations (Fujii, H. et al., 2016; Guilliams, M. et al., 2016). However, current methods are limited by ambiguities in differentiating true signals and autofluorescence, which hampers the full adaptation to high-dimensional analysis of expression of many proteins (Brummelman, J. et al., 2019; Lokwani, R. et al., 2022).
[0007] Recent advancements have introduced the use of temporal data, such as the time elapsed since a cellular event, to provide a more dynamic view of cellular behaviour. These methods include pseudotime analysis (Tan, B.J.Y. et al., 2021) and RNA velocity analysis (La Manno, G. et al., 2018), which usually use high-throughput sequencing of transcripts (gene expression analysis), and the Timer-of-cell-kinetics-and-activity (Tocky), utilizing a Fluorescent Timer protein that spontaneously changes its emission spectrum with known kinetics. However, integrating temporal data with multidimensional protein expression profiles to elucidate cell lineage dynamics remains a challenge.
[0008] The present invention overcomes these limitations by introducing a novel method that employs Canonical Correspondence Analysis (CCA), enabling the analysis of time-linked data and multidimensional gene (protein) expression profiles. Various variant methods, including Redundancy Analysis (RDA), have been developed. Together with CCA, these methods are classified as Constrained Ordination techniques, sharing key mathematical analysis steps, albeit with variations in the weighting of intermediate matrices and outputs.
[0009] The work underlying the present invention was supported by the Medical Research Council [grant number MR / S000208 / 1]; and the Biotechnology and Biological Sciences Research Council [grant number BB / J013951 / 2],
[0010] Summary of the invention
[0011] The present inventors have investigated an approach for identifying temporally distinct cell lineages following a cellular event such as a cell signalling event. The approach canonically combines data indicating time elapsed following a cellular event such as a cell signalling event, with expression data for marker proteins, using a novel analysis method, to group cells according to their cellular activation and differentiation status. The method allows identification of unique cell lineages within a cell population by analysing how the developing cells change marker protein expression profiles over time, rather than simply by expression profile or time alone. Collectively, the methods of the present invention provide means to analyse cellular trajectories and their protein expression changes across time during activation, differentiation, or development, enabling identification of cell lineage and functional status with a level of precision and detail not achievable with existing methods.
[0012] Accordingly, in a first aspect the present invention provides a computer-implemented method comprising the following steps:
[0013] (i) accessing time-linked expression data for a specific protein in a plurality of cells, wherein expression of the specific protein is linked to a cellular event; and processing the time-linked expression data to temporally classify the plurality of cells according to time elapsed following the cellular event, thereby obtaining temporally classified data;
[0014] (ii) accessing expression data for a plurality of proteins expressed by the plurality of cells of step (i); and processing the expression data for the plurality of proteins to categorise the plurality of cells into clusters of cells according to the relative levels of expression of the plurality of proteins; and,
[0015] (iii) processing the temporally classified data and the clusters of cells in combination to group the clusters of cells according to their cellular activation and differentiation status.
[0016] Detailed description of the invention
[0017] In a first aspect, the present invention relates to a computer-implemented method comprising three distinct steps (i)-(iii). The method of the first aspect is suitable for, but not limited to, identifying temporally distinct cell lineages following a cellular event such as a cell signalling event. The method of the first aspect is also suitable for, but not limited to, grouping cells according to their cellular activation and differentiation status. The method involves accessing and processing two sets of data relating to a plurality of cells, wherein the first set of data provides a measure of time elapsed following a cellular event, and the second set of data provides a measure of the level of expression of a plurality of proteins by the plurality of cells. The method of the first aspect provides a way of analysing both data sets together to provide novel and surprising insights into cell lineage and status.
[0018] The three distinct steps of the method are numbered (i) to (iii) in the following application. The steps, in particular steps (i) and (ii), are not limited to being performed sequentially or in any specific order, and combinations of steps can be performed simultaneously or in a different order to the numbering provided herein. Nevertheless, step (iii) relies on the temporally classified data and the clusters of cells identified in steps (i) and (ii) respectively.
[0019] Step (i)
[0020] Step (i) of the method of the first aspect of the invention is accessing time-linked expression data for a specific protein in a plurality of cells, wherein expression of the specific protein is linked to a cellular event, and processing the time-linked expression data to temporally classify the plurality of cells according to time elapsed following the cellular event, thereby obtaining temporally classified data.
[0021] By “accessing” is meant retrieving or obtaining information or data from a storage medium or a data source. The data source can be any repository where data is stored, including but not limited to, a dataset, a database, or a server.
[0022] By “time-linked expression data” is meant expression data that indicates time elapsed since a given point in time. The time-linked expression data is generated using protein expression. The data comprises measurements of protein expression over time. Typically, the data is in the form of measurements of fluorescence emitted by an expressed construct over time, wherein the measurements detail at least the colour and intensity of fluorescence. The data therefore links the protein expression to the time elapsed.
[0023] By “cellular event” is meant an event which occurs at a molecular level in or on the surface of a cell, and which triggers the expression of one or more proteins. The cellular event is typically a cell signalling event, but can also be the onset of expression of any protein of interest for example in response to internal or external stimuli. The stimuli can be binding between a receptor and a ligand, such as binding between a T-cell receptor and an antigen.
[0024] Typically, the specific protein is a protein whose expression is activated concurrently with the occurrence of the cellular event being investigated. Accordingly, the expression of the specific protein is linked to the cellular event. Therefore, the time-linked expression data, which indicates the amount of time that the specific protein has been expressed for, also indicates the amount of time elapsed since the occurrence of the cellular event.
[0025] The method of the first aspect of the present invention can comprise generating the time- linked expression data, or alternatively can be limited to accessing pre-existing time-linked expression data. The following aspects of the invention related to generating time-linked expression data can be included in the methods of the first aspect of the invention as additional steps prior to accessing the time-linked expression data.
[0026] The time-linked expression data has typically been generated by expressing the specific protein using a construct that additionally encodes a fluorescent protein. Fluorescent proteins for use in the invention can be mutants of mCherry. Fluorescent proteins for use in the invention can be the result of combination of two fluorescent proteins with different expression dynamics into a single fusion protein. Fluorescent proteins for use in the invention include fastFT, dsRedE5TIMER and TandemFP. TandemFP is a fusion protein consisting of DsRed2 and sfGFP. The construct is typically a genetic construct, or an engineered assembly of genetic material in sequence, where the gene corresponding to the specific protein is combined with the gene expressing the fluorescent protein in sequence such that expression of both the specific protein and the fluorescent protein are activated and driven together. This can be achieved by having a single promoter driving expression of both proteins. When a cellular event causes expression of the specific protein to begin, expression of the fluorescent protein begins at the same time, and therefore fluorescence can be measured following the cellular event. The construct can be an Nr4a3-Tocky DNA construct as described in Bending et al. (Bending, et al., 2018). .
[0027] Typically, the colour of fluorescence of the fluorescent protein changes over time. Fluorescent proteins whose colour changes over time are known in the art and include fastFT, dsRedE5TIMER and TandemFP (Subach, et al., 2009) (Yau, et al., 2020) (Khmelinskii, et al., 2012). The change in colour of the fluorescence from a first colour to a second colour can be such that the first colour of fluorescence has a first half-life time, whilst the second colour of fluorescence has a second half-life time. The first half-life time indicates the amount of time it takes for half of the fluorescence of a plurality of fluorescent proteins to change from the first colour to the second colour. The second halflife time indicates the amount of time it takes for half of the plurality of fluorescent proteins to decay such that they no longer emit any fluorescence. Where the fluorescent protein is fastFT, the first half-life time can be 4 hours, and the second half-life time can be 122 hours. The fluorescent protein can be a mutant of mCherry, such as fastFT as described in Subach et al. (Subach, et al., 2009).
[0028] Where the first colour of fluorescence has a first half-life time, and the second colour of fluorescence has a second half-life time, the ratio of one colour of fluorescence compared to the other colour of fluorescence can be used as a measure of time elapsed since expression of the protein began. Where the fluorescent protein is a non-tandem protein, such as fastFT or dsRedE5TIMER, this is because each time a new fluorescent protein is transcribed, it will emit the first colour of fluorescence until a given time point, at which point the colour changes, and over time as more individual proteins reach the colour change time point the overall fluorescence emitted will shift from the first colour to the second colour. Where the fluorescent protein is a tandem protein, such as TandemFP, the change of overall fluorescence colour emitted is caused by a difference between the maturation kinetics between two fluorescent proteins, such as sfGFP and DsRed2, that means as one fluorescent protein degrades quicker than the other, the overall colour shifts towards that of the slower degrading protein. This allows linking of expression of the specific protein to time, via the fluorescence of the fluorescent protein.
[0029] Where time-linked expression data has been generated by expressing a specific protein using a construct that additionally encodes a fluorescent protein, the time-linked expression data comprises fluorescence measurements, in particular the time-linked expression data comprises measurements of fluorescence over time.
[0030] Time-linked expression data is not limited to changes in fluorescence colour. Time-linked expression data can also be generated using pseudotime analysis wherein the data represents inferences of time elapsed obtained via similarity analysis of expression profiles over time. Time-linked expression data can also be generated using RNA velocity methods wherein the data represents estimations of changes in gene expression, relying on decay rates of mRNA molecules to provide a measure of time elapsed.
[0031] Typically, the time-linked expression data has been generated using flow cytometry.
[0032] The specific protein is typically a marker protein, for example Nr4a3, Foxp3, IL2, or PD1.
[0033] By “plurality of cells” is meant at least two cells, for example at least 100 cells, at least 500 cells, at least 1000 cells, at least 5000 cells, at least 10,000 cells, at least 50,000 cells, at least 100,000 cells, at least 500,000 cells, or at least 1,000,000 cells.
[0034] The cells of the plurality of cells can be T-cells, for example CD4+ T-cells, CD8+ T-cells, regulatory T-cells, effector T-cells, memory T-cells, natural killer T-cells, or gamma delta T-cells. The CD4+ T-cells can be Th1 T-cells, Th2 T-cells, Th17 T-cells, or T follicular helper cells. The regulatory T-cells can be native regulatory T-cells, effector regulatory T- cells, or T follicular regulatory cells. The memory T-cells can be central memory T-cells, effector memory-T cells, or tissue resident memory T-cells. The CD4+ or CD8+ T-cells can be naive T-cells. A naive T-cell is a T-cell that has not yet encountered and been activated by a specific antigen, and therefore has not differentiated into their final state yet.
[0035] The cellular event can be T-cell receptor (TCR) signalling. T-cell receptor signalling, or T- cell receptor activation, is the initiation of a cascade of biochemical pathways following the binding of a specific antigen to a T-cell receptor on the surface of a T-cell. The resulting cascade drives T-cell development, selection and immune response.
[0036] The cells of the plurality of cells can be B-cells, for example plasma cells, memory B-cells, pre-B cells, mature B-cells, follicular B-cells, marginal zone B-cells, transitional B-cells, or B1 cells.
[0037] The cellular event can be B-cell receptor signalling. B-cell receptor signalling, or B-cell receptor activation, is the initiation of a cascade of biochemical pathways following the binding of a specific antigen to a B-cell receptor on the surface of a B-cell. The resulting cascade drives B-cell development, selection and immune response.
[0038] The cells of the plurality of cells can alternatively be innate lymphoid cells or dendritic cells.
[0039] More broadly, the method of the first aspect of the present invention is applicable to any cellular event that relates to the onset of expression of a protein. The cellular event can be the onset of expression of a marker protein, for example a protein selected from the group consisting of Nr4a3, Foxp3, IL2, and PD1.
[0040] By “processing” is meant the manipulation, transformation, or analysis of data to derive insights, to generate desired outputs, or facilitate specific functionalities. Processing data can include the application of algorithms, statistical methods, or computational techniques to extract patterns, trends, or information from data. Computational techniques can include performing calculations or mathematical operations on data. Processing data can also include managing the storage of data, including data retrieval and data organization. Processing data can also include transmission or transfer of data between systems, devices or components. Accordingly, “processing the time-linked expression data” in step (i) encompasses manipulating the time-linked expression data for each cell so as to identify the time elapsed since the cellular event occurred in or on the cell, and subsequently classifying the cells into groups according to time elapsed following the cellular event. The resulting groups of classified cells are referred to herein as temporally classified data.
[0041] Processing the time-linked expression data can comprise the use of trigonometric transformation. For example, trigonometric transformation can be applied to the fluorescence data to arrive at an angle and an intensity for each cell, where the angle value of each cell indicates time elapsed following a cellular event such as a cell signalling event, and can be classified into a plurality of classifications defined by ranges of angle values. In such embodiments, a timer angle of 0° typically indicates that a cell has undergone the cellular event very recently, whereas a timer angle of 90° indicates that a cell has been removed from the cellular event and terminated transcription of the specific protein that is linked to the cellular event. The timer angle can also indicate frequency of transcription, in particular between 45° and 90°, as the angle approaches 45° it indicates an increasing frequency of transcription, and as the value approaches 90° it indicates a decreasing frequency of transcription. The intensity value of each cell indicates the strength of the biochemical signals initiated by the signalling event. The classifications defined by ranges of angle values can be “new”, “persistent”, and “arrested”, wherein each classification title refers to the state of the post-cellular event expression of the specific protein, or the state of post-cellular event signalling. Accordingly, the temporally classified data can use classifications comprising: newly started post-event expression, persistent post-event expression and arrested post-event expression, wherein the expression is expression of the specific protein.
[0042] The output of step (i) of the method of the first aspect of the invention is temporally classified data. The temporally classified data can therefore comprise data that classifies the plurality of cells according to time elapsed, and the state of expression of the specific protein, following a cellular event such as a cell signalling event.
[0043] Step (ii)
[0044] Step (ii) of the method of the first aspect of the invention is accessing expression data for a plurality of proteins expressed by the plurality of cells of step (i), and processing the expression data for the plurality of proteins to categorise the plurality of cells into clusters of cells according to the relative levels of expression of the plurality of proteins. “Accessing” and “processing” the data in step (ii) has the same general meaning as in step (i), albeit processing the data in step (ii) typically involves different mathematical and computational methods, as set out below. The plurality of cells is the same plurality of cells as defined in relation to step (i).
[0045] The method of the first aspect of the present invention can comprise generating the expression data for a plurality of proteins, or alternatively can be limited to accessing preexisting expression data for a plurality of proteins. The following aspects of the invention related to generating expression data for a plurality of proteins can be included in the methods of the first aspect of the invention as additional steps prior to accessing the expression data for a plurality of proteins.
[0046] Cells at different stages of differentiation or cell development have differing expression profiles due to the relative differences in expression levels of certain proteins that change over the differentiation or development process. The plurality of proteins in step (ii) is a plurality of proteins selected for their ability to provide insight into the status, for example the differentiation status or the development status, of the cell. These can be proteins that are known to change in expression level over the course of cellular differentiation or development. The plurality of proteins in step (ii) is typically a plurality of marker proteins. For example, the plurality of proteins in step (ii) may include one or more proteins selected from the group consisting of Foxp3, CD25, PD-1, CTLA-4, CD69, OX-40, ICOS, GITR, TNFR2, 4-1 BB, CD44, CD45RB, CD122, LAG3, TIGIT, CD4, and CD8,. The plurality of proteins in step (ii) can include the specific protein from step (i).
[0047] The expression data for a plurality of proteins can be generated using one of flow cytometry, RNA sequencing, microscopy (such as confocal microscopy) and T-cell receptor sequencing. The expression data provides information on the level of expression of each of the plurality of proteins.
[0048] Where flow cytometry is used to generate the expression data used in step (ii), the plurality of proteins typically comprises 5 or more proteins, 10 or more proteins, 25 or more proteins, or up to 50 proteins, or preferably between 10 and 20 proteins. Where RNA sequencing is used to generate the expression data used in step (ii), the plurality of proteins can comprise 50 or more proteins, 100 or more proteins, 250 or more proteins, or 500 or more proteins. The expression data for a plurality of proteins is processed using any suitable method known in the art, for example principle component analysis (PCA), autofluorescence identification, and / or clustering. Where flow cytometry has been used to generate the expression data for a plurality of proteins, autofluorescence identification, and determination of the range of autofluorescence, can be used to define expression of each of the plurality of proteins. In particular, the expression data can be normalised by defining autofluorescence and determining a threshold for true expression signals above autofluorescence for each marker protein, thereby reducing noise in the expression data. PCA can be applied to the expression data and the resulting principle component scores can be used for clustering. The clustering can be carried out using any suitable clustering method, including k-means clustering, self-organising map clustering, hierarchical clustering, and model based clustering. In these embodiments, the processing results in a fragmented PCA plot which allows identification of distinct cell clusters according to the relative levels of expression of the plurality of proteins (their expression profiles), and therefore their differentiation status, or development status.
[0049] The output of step (ii) of the method of the first aspect of the invention is that the plurality of cells are categorised into clusters of cells according to the relative levels of expression of the plurality of proteins. By “clusters of cells” is meant groups of cells wherein each group of cells shares similarities in their expression profiles. In particular, each group of cells will have levels of expression of given proteins that are similar relative to cells in other groups. The groups of cells differ in that some groups of cells will express given proteins more or less than others, reflecting differences in differentiation status, or development status.
[0050] Step (Hi)
[0051] Step (iii) of the method of the first aspect of the invention is processing the temporally classified data (obtained in step (i)) and the clusters of cells (obtained in step (ii)) in combination to group the clusters of cells according to their cellular activation and differentiation status.
[0052] “Processing” the data in step (iii) has the same general meaning as in steps (i) and (ii), albeit processing the data in step (iii) typically involves different mathematical and computational methods, as set out below. The processing in step (iii) is typically carried out using a method comprising a constrained ordination technique. Typically the constrained ordination technique used is canonical correspondence analysis (CCA). CCA can be used to project the clustered expression data onto a subspace defined by the temporally classified data, calibrating similarities in expression profile between single cells given the progressive increase of angle or intensity values in the temporally classified data. The output of the CCA can include the mean position, or barycentre, for each cluster within the subspace defined by the temporally classified data. This mean represents the average similarity score for the cells within each cluster, reflecting both their expression profiles and their angle and intensity values. The constrained ordination technique can also be redundancy analysis.
[0053] The processing in step (iii) can be carried out using a method additionally comprising network analysis. Network analysis can comprise treating each cluster barycentre as a node, and defining edges between the nodes based on the similarity scores between clusters, with each edge assigned a weight corresponding to this similarity. A threshold can be applied to these weights to determine which edges are significant enough to be included in the network visualization, ensuring that only meaningful relationships between clusters are represented. This analysis can result in ordering the clusters of cells according to their mean angle values, to determine the temporal sequence of clusters in differentiation or development. Such analysis allows for the clusters to be identified as consisting of immature and early stage cells that have just received signalling, or cells that are more mature and later in development following signalling, or cells in between those states. Network analysis can comprise creating a network graph based on the expression profile similarity between the clusters of step (ii) and the cell categories of the temporally classified data (i.e. newly started post-event expression, persistent post-event expression and arrested post-event expression). The resulting network graph groups the clusters of cells according to their cellular activation and differentiation status, or in other words, it allows the identification of temporally distinct lineages which group cells according to their expression profiles over time following the signalling event.
[0054] The processing in step (iii) typically produces a network of cell clusters connected according to similarities in their cellular activation and differentiation status.
[0055] The processing in step (iii) therefore allows cross-analysis of the clustered expression data and the temporally classified data to reveal multidimensional cell-cell relationships and the progression of differentiation or development. The multidimensional analyses of step (iii) provides both new and surprising insights and ensures that the clustering is accurate and comprehensive, which is important to ensure the right clusters of cells are identified in applications of the method. In particular, the analyses can map the evolution of cell lineages over time, tracking transitions through different states of activation and differentiation, and characterize the complex patterns of protein expression that define cellular functions and identities across these lineages. There is currently no other method for achieving such canonical analysis.
[0056] Further aspects of the invention
[0057] The methods of the present invention can be carried out on systems or using computer readable storage mediums containing instructions for carrying out the methods.
[0058] Figure 9 illustrates a block diagram of one implementation of a computing system 900 within which a set of instructions, for causing the computing system to perform any one or more of the methodologies discussed herein, may be executed. In implementations, the computing system may be connected (e.g., networked) to other machines in a Local Area Network (LAN), an intranet, an extranet, or the Internet. The computing system may operate in the capacity of a server or a client machine in a client-server network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The computing system may be a personal computer (PC), a tablet computer, a set-top box (STB), a Personal Digital Assistant (PDA), a cellular telephone, a web appliance, a server, a network router, switch or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by that machine. Further, while only a single computing system is illustrated, the term “computing system” shall also be taken to include any collection of machines (e.g., computers) that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein.
[0059] The example computing system 900 includes a processing device 902, a main memory 904 (for example, read-only memory (ROM), flash memory, or dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM)), a static memory 906 (for example, flash memory or static random access memory (SRAM)), and a secondary memory (for example, a data storage device 918), which communicate with each other via a bus 930.
[0060] Processing device 902 represents one or more general-purpose processors such as a microprocessor, central processing unit, or the like. Processing device 902 may also be one or more special-purpose processing devices such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), network processor, or the like. Processing device 902 is configured to execute the processing logic (instructions 922) for performing the operations and steps discussed herein.
[0061] The computing system 900 may further include a network interface device 908. The computing system 900 also may include a video display unit 910 (for example, a liquid crystal display (LCD) or a cathode ray tube (CRT)), an alphanumeric input device 912 (for example, a keyboard or touchscreen), and a cursor control device 914 (for example, a mouse or touchscreen).
[0062] The data storage device 918 may include one or more machine-readable storage media (or more specifically one or more non-transitory computer-readable storage media) 928 on which is stored one or more sets of instructions 922 embodying any one or more of the methodologies or functions described herein. The instructions 922 may also reside, completely or at least partially, within the main memory 904 and / or within the processing device 902 during execution thereof by the computer system 900, the main memory 904 and the processing device 902 also constituting computer-readable storage media.
[0063] The various methods described above may also be implemented by a computer program. The computer program may include computer code (e.g. instructions) 1010 arranged to instruct a computer to perform the functions of one or more of the various methods described above. The steps of the methods described above may be performed in any suitable order. For example, step (i) of the method of the first aspect of the invention may be performed before, after, simultaneously or substantially simultaneously with step (ii). The computer program and / or the code 1010 for performing such methods may be provided to an apparatus, such as a computer, on one or more computer readable media or, more generally, a computer program product 1000)), depicted in Figure 10. The computer readable media may be transitory or non-transitory. The one or more computer readable media 1000 could be, for example, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, or a propagation medium for data transmission, for example for downloading the code over the Internet. Alternatively, the one or more computer readable media could take the form of one or more physical computer readable media such as semiconductor or solid state memory, magnetic tape, a removable computer diskette, a random access memory (RAM), a read-only memory (ROM), a rigid magnetic disc, and an optical disk, such as a CD-ROM, CD-R / W or DVD. The instructions 1010 may also reside, completely or at least partially, within the memory 913 and / or within the controller circuitry 911 during execution thereof by the computing system 910, the memory 913 and the controller circuitry 911 also constituting computer- readable storage media.
[0064] Accordingly, in a second aspect, the present invention relates to a computer readable storage medium containing instructions which, when executed on one or more data processors, cause the one or more data processors to perform the method of the first aspect of the invention. The computer readable storage medium can be transitory or non- transitory.
[0065] Furthermore, in a third aspect, the present invention relates to a system comprising a memory and one or more processors, wherein the one or more processors is configured to perform the method of the first aspect of the invention.
[0066] Preferred features for the second and subsequent aspects of the invention are as for the first aspect mutatis mutandis. It will be appreciated that all embodiments described herein are considered to be broadly applicable and combinable with any and all other consistent embodiments, as appropriate. Such combinations are considered to fall within the scope of the present invention.
[0067] The present invention will now be further described by way of reference to the following Example which is present for the purposes of illustration only. In the Example, reference is made to a number of Figures in which:
[0068] Figure 1: A schematic illustration of the experimental setup to analyse the effects of anti- PD-L1 antibody treatment on developing thymocytes. Flow cytometric analysis of developing T-cells was performed on thymi harvested from Nr4a3-Tocky mice on day 4 following the administration of the anti-PD-L1 antibody on days 0 and 2.
[0069] Figure 2: Following autofluorescence identification, Principal Component Analysis (PCA) was applied to the flow cytometric data to reduce dimensionality and identify distinct T-cell clusters based on their marker expression profiles. Left: Cells are plotted with colour codes representing Tocky loci. Middle: PCA clusters are highlighted and differentiated by colour codes. Right: Cluster identities are indicated by cluster numbers, shown at the barycentre of each cluster. Figure 3: Dot plot displaying the percentage of marker expression and expression levels, represented by dot size and heat colour coding, respectively.
[0070] Figure 4: This bar plot illustrates the percentage of cells within each cluster, comparing the two groups. Statistical significance is indicated by an asterisk (*p < 0.05), determined by the Mann-Whitney test with p-value adjustment.
[0071] Figure 5: Biplot of Canonical Correspondence Analysis (CCA) of T-cell Development Markers and Timer Angle / lntensity, showing developing T-cells by dots and explanatory variables by arrows (Timer-Angle and -Intensity). Colours indicate the levels of Timer- Angle (left) or Timer-Intensity (right) as shown in the colour code.
[0072] Figure 6: A CCA network diagram is constructed using PCA clusters as nodes within the CCA space. Edges are determined by the distances between cluster barycentres.
[0073] Figure 7: T-cell developmental trajectories within the CCA network are illustrated, with identified trajectories highlighted by bold pink edges and nodes. Upper left: The trajectory from cluster C14 to C10, representing maturation from double-positive to CD4-single positive. Upper right: The trajectory from C2 to C15, depicting differentiation from doublepositive to CD8-single positive. Lower: The trajectory from clusters C14 / 7 / 2 to C8, indicating a distinct pathway of T-cell development.
[0074] Figure 8: Line graphs depicting the percentage of cells in each Tocky locus and their marker expression along defined Timer trajectories. Statistical significance is marked by an asterisk (*p < 0.05), as determined by the Mann-Whitney test with p-value adjustment.
[0075] Figure 9: Block diagram of one implementation of a computing system 900 within which a set of instructions, for causing the computing system to perform any one or more of the methodologies discussed herein, may be executed.
[0076] Figure 10: Block diagram of one implementation of a computer program product 1000 for performing any one or more of the methodologies discussed herein.
[0077] Example
[0078] Application of the Computational Algorithms to Elucidate Anti-PD-L1 Therapy Effects on Thymic T-Cell Development Given that PD-1 expression is induced in developing T-cells in the thymus, we explored whether cancer immunotherapy through PD-1 blockade influences T-cell development. We utilized an anti-PD-L1 antibody that binds to the PD-1 ligand on immune cells, thereby inhibiting PD-1 signalling. Nr4a3-Tocky mice were administered the anti-PD-L1 antibody on days 0 and 2. Subsequently, thymocytes were analysed on day 4 by flow cytometry (Figure 1).
[0079] Principal component analysis (PCA) was applied to the flow cytometric data after identifying autofluorescence thresholds. K-means clustering was used to identify T-cell clusters with distinct marker profiles in the PCA space (referred to as PCA clusters, Figures 2 and 3). Clusters 1 , 3, 8, and 16 increased following anti-PD-L1 therapy, while clusters 4 and 15 decreased (Figure 4). Canonical Correspondence Analysis (CCA) was then employed to analyze the progression of Timer-Angle and Timer-Intensity across multidimensional marker profiles (Figure 5).
[0080] A network was constructed using PCA clusters as nodes within the CCA space. Utilizing CCA's capability for similarity analysis through its metrics, the network edges were determined by the distance between the barycenters of the PCA clusters. Figure 6 illustrates the changes in T-cell percentages within clusters (Iog2 fold change), Tocky- Angle, and -Intensity, and the expression of CD69 and PD-1. T-cells develop from the bottom of the network (low Timer Angle and high CD69 expression) to the top (high Timer Angle and low CD69 expression).
[0081] The Timer trajectory of developing thymic T cells is defined by assuming a unidirectional increase of Timer-Angle values, choosing the closest paths in the CCA space. The trajectory from the immature double-positive (DP) cluster C14 to the mature CD4-single positive (SP) cluster C10 was unambiguously determined (Figure 7), capturing the progressive decrease of CD8 expression and the induction of PD-1, GITR, and Foxp3 in sequence (Figure 8). The trajectory from the immature cluster C2 to the mature CD8-SP cluster C15 identified the progressive differentiation of DP cells into CD8-SP cells, the persistent decline of CD4 expression, and the induction of PD-1 and GITR, but not Foxp3, as is typically seen in CD8-SP lineage cells.
[0082] Moreover, it was noted that the immature cluster C8 receives trajectories from all the most immature DP cell clusters, C2, C7, C11, and C14, with remarkably high PD-1 expression. Interestingly, the trajectories C14-C10 and C2-C15 accumulated cells at the Arrested locus, indicating that these trajectories included a significant portion of T-cells that develop and survive as they mature Timer proteins. Conversely, the trajectory C14-C8 barely contained any cells apart from at the New locus, suggesting that the majority of cells do not reside within the trajectory by the time the Timer matures to become Arrested cells.
[0083] The lack of significant differences in cell number kinetics across the three trajectories suggests that cell kinetics post-TCR signalling were not significantly altered by the therapy. While the majority of markers did not show any significant differences by the treatment, PD-1 expression exhibited interesting changes between the two groups. T-cells in the trajectory C14-C10 upregulated PD-1 expression throughout the Timer loci, in contrast to the significant increase in PD-1 expression only in the Timer-negative cells and at the New locus in the trajectories C2-C15 and C14-C8. Notably, PD-1 expression was exceptionally high in the trajectory C14-C8, especially at the NPt locus.
[0084] As PD-1 signalling suppresses PD-1 expression itself, the results support the notion that PD-1 signalling conveys negative signals in the three trajectories, inducing PD-1 high thymocytes that accumulated Timer proteins and showed high Timer Intensity levels, remarkably in the clusters C7, C8, C3, and C1 (Figure 6). It is also likely that PD-1 very high thymocytes in the trajectory C14-C8 succumbed to negative selection and were removed from the system.
[0085] References
[0086] Holguera, I. & Desplan, C. Neuronal specification in space and time. Science 362, 176- 180 (2018).
[0087] Raimundo, J. & Levine, M. How do genomes encode developmental time? Genes Dev 37, 37-40 (2023).
[0088] Voortman, L. et al. Temporally dynamic antagonism between transcription and chromatin compaction controls stochastic photoreceptor specification in flies. Dev Cell 57, 1817- 1832.e1815 (2022).
[0089] Fujii, H. et al. Regulatory T cells in melanoma revisited by a computational clustering of FOXP3+ T cell subpopulations. The Journal of Immunology 196, 2885-2892 (2016). Guilliams, M. et al. Unsupervised High-Dimensional Analysis Aligns Dendritic Cells across Tissues and Species. Immunity 45, 669-684 (2016).
[0090] Brummelman, J. et al. Development, application and computational analysis of highdimensional fluorescent antibody panels for single-cell flow cytometry. Nature Protocols 14, 1946-1969 (2019). Lokwani, R., Chaudhari, R., Wolf, M.T. & Sadtler, K. Spectral cytometry on highly autofl uorescent samples. Nature Reviews Methods Primers 2, 71 (2022).
[0091] Tan, B.J.Y. et al. HTLV-1 infection promotes excessive T cell activation and transformation into adult T cell leukemia / lymphoma. The Journal of Clinical Investigation 131 (2021).
[0092] La Manno, G. et al. RNA velocity of single cells. Nature 560, 494-498 (2018).
[0093] Bending, D. et al. A timer for analyzing temporally dynamic changes in transcription during differentiation in vivo. J Cell Biol 217 (8), 2931-2950 (2018).
[0094] Khmelinskii, A. et al. Tandem fluorescent protein timers for in vivo analysis of protein dynamics. Nature Biotechnology 30, 708-714 (2012).
[0095] Subach, F. et al. Monomeric fluorescent timers that change color from blue to red report on cellular trafficking. Nat Chem Biol. 5 (2), 118-126 (2009).
[0096] Yau, B. et al. A fluorescent timer reporter enables sorting of insulin secretory granules by age. J. Biol. Chem. 295 (27), 8901-8911 (2020).
Claims
CLAIMS1) A computer-implemented method comprising the following steps:(i) accessing time-linked expression data for a specific protein in a plurality of cells, wherein expression of the specific protein is linked to a cellular event; and processing the time-linked expression data to temporally classify the plurality of cells according to time elapsed following the cellular event, thereby obtaining temporally classified data;(ii) accessing expression data for a plurality of proteins expressed by the plurality of cells of step (i); and processing the expression data for the plurality of proteins to categorise the plurality of cells into clusters of cells according to the relative levels of expression of the plurality of proteins; and,(iii) processing the temporally classified data and the clusters of cells in combination to group the clusters of cells according to their cellular activation and differentiation status.2) The computer-implemented method of claim 1, wherein the time-linked expression data has been generated by expressing the specific protein using a construct that additionally encodes a fluorescent protein.3) The computer-implemented method of claim 2, wherein the colour of fluorescence of the fluorescent protein changes over time.4) The computer-implemented method of any of claims 2-3, wherein the time-linked expression data comprises fluorescence measurements.5) The computer-implemented method of any preceding claim, wherein the time- linked expression data has been generated using flow cytometry.6) The computer-implemented method of any preceding claim, wherein the specific protein of step (i) is a marker protein.7) The computer-implemented method of claim 6, wherein the specific protein of step (i) is selected from the group consisting of Nr4a3, Foxp3, IL2, and PD1.8) The computer-implemented method of any preceding claim, wherein the cellular event is onset of expression of a protein, optionally wherein the cellular event is the onset of expression of a protein selected from the group consisting of Nr4a3, Foxp3, IL2, and PD1.9) The computer-implemented method of any preceding claim, wherein the plurality of cells are T-cells; optionally wherein the T-cells are CD4+ T-cells, CD8+ T-cells, regulatory T- cells, effector T-cells, memory T-cells, natural killer T-cells, or gamma delta T-cells; and, further optionally wherein the CD4+ T-cells are Th1 T-cells, Th2 T- cells, Th17 T-cells, or T follicular helper cells; wherein the regulatory T-cells are native regulatory T-cells, effector regulatory T-cells, or T follicular regulatory cells; wherein the memory T-cells are central memory T-cells, effector memory-T cells, or tissue resident memory T-cells; or, wherein the CD4+ or CD8+ T-cells are naive T-cells.10) The computer-implemented method of claim 9, wherein the cellular event is T-cell receptor signalling.11) The computer-implemented method of any one of claims 1 to 7, wherein the plurality of cells are B-cells; optionally wherein, the B-cells are plasma cells, memory B-cells, pre-B cells, mature B-cells, follicular B-cells, marginal zone B-cells, transitional B- cells, or B1 cells; further optionally wherein the cellular event is B-cell receptor signalling.12) The computer-implemented method of any one of claims 1 to 7, wherein the plurality of cells are innate lymphoid cells or dendritic cells.13) The computer-implemented method of any preceding claim, wherein processing the time-linked expression data in step (i) comprises the use of trigonometric transformation.14) The computer-implemented method of any preceding claim, wherein the temporally classified data uses classifications comprising: newly started post-event expression, persistent post-event expression and arrested post-event expression.15) The computer-implemented method of any preceding claim, wherein the plurality of proteins in step (ii) is a plurality of marker proteins.16) The computer-implemented method of any preceding claim, wherein the plurality of proteins in step (ii) includes one or more proteins selected from the group consisting of Foxp3, CD25, PD-1, CTLA-4, CD69, OX-40, ICOS, GITR, TNFR2, 4- 1BB, CD44, CD45RB, CD122, LAG3, TIGIT, CD4, and CD8.17) The computer-implemented method of any preceding claim, wherein the plurality of proteins in step (ii) includes the specific protein from step (i).18) The computer-implemented method of any preceding claim, wherein processing in step (ii) is carried out using a method comprising at least one of principle component analysis (PCA), autofluorescence identification, and / or k-means clustering.19) The computer-implemented method of any preceding claim, wherein the expression data in step (ii) has been generated using one of flow cytometry, RNA sequencing and T-cell receptor sequencing.20) The computer-implemented method of any preceding claim, wherein the processing in step (iii) is carried out using a method comprising canonical correspondence analysis (CCA), optionally wherein the method further comprises network analysis.21) The computer-implemented method of any preceding claim, wherein the plurality of proteins is 5 or more proteins, 10 or more proteins, 50 or more proteins, 100 or more proteins, or 500 or more proteins. 22) A computer readable storage medium containing instructions which, when executed on one or more data processors, cause the one or more data processors to perform the method of any preceding claim.23) A system comprising a memory and one or more processors, wherein the one or more processors is configured to perform the method of any one of claims 1 to 21.