Method of generating an input image for a classifier for classifying a sample of cells
A novel method converts fluorescence measurements into a 2D input image for classification, addressing the challenge of analyzing smaller datasets in single-cell gene expression dynamics, enabling precise cell classification and transcription insights.
Patent Information
- Application Number
- PCT/EP2025/072002
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-31
- Filing Date
- 2025-07-30
- Publication Date
- 2026-02-05
AI Technical Summary
Existing methods for analyzing single-cell gene expression dynamics using fluorescent timer proteins face challenges in accurately capturing and measuring temporal elements, particularly in smaller datasets, and often require large sample sizes and extensive image data, which is not suitable for certain flow cytometric studies.
A novel method is developed to convert fluorescence measurements into a 2D input image for classification using a classifier, enabling precise analysis of cell samples by determining intensity and mixing ratios of composite fluorescence colors, which can be processed by image classifiers like convolutional neural networks.
This approach allows for accurate classification of cell samples, providing unique insights into transcription mechanisms and overcoming the limitations of traditional strategies, especially in smaller datasets.
Smart Images

Figure EP2025072002_05022026_PF_FP_ABST
Abstract
Description
[0001] METHOD OF GENERATING AN INPUT IMAGE FOR A CLASSIFIER FOR CLASSIFYING A SAMPLE OF CELLS
[0002] Field of the invention
[0003] The present invention relates to a method of generating an input image for a classifier for classifying a sample of cells as one of a plurality of cell classes, and related methods of classifying the sample of cells and of training the classifier, as well as computer-readable storage media and systems for implementing the methods.
[0004] Background to the invention
[0005] Understanding the temporal dynamics of gene expression is fundamental to comprehending gene regulation's functional significance, cellular differentiation, and development across various disciplines in molecular, cellular, and systems biology. Single-cell level analyses of diverse tissues have utilized trajectory analyses to reconstruct in vivo activation or differentiation processes. However, trajectory analysis inherently relies on similarity analysis through dimensional reduction, which is not a direct measurement of temporal dynamics in vivo. Thus, a challenge lies in accurately capturing and measuring temporal elements of these processes within individual cells.
[0006] Fluorescent Timer proteins offer a unique solution to this challenge. Previous approaches, such as mathematical modelling of Fluorescent Timer expression using tandem timers, have provided insights into cell population dynamics1. However, the integration of singlecell techniques with a Timer-based approach has yet to be realized. The Timer-of-cell- kinetics-and-activity (Tocky) system for single-cell flow cytometric analysis of cellular activities and transcription2was developed previously. Tocky uses a fluorescent timer protein, which spontaneously and irreversibly changes its chromophore from one colour to another over a set time period3.
[0007] Given the complex kinetics of involving gene expression through transcription and cellular dynamics, it is essential to establish a data-oriented method for quantitatively analysing Timer fluorescence at the single-cell level. ML approaches have been successfully applied to multidimensional flow cytometric data in various clinical studies. For example, a combination of FlowSOM and Random Forest (RF) has been employed to analyse clinical data4, while Convolutional Neural Networks (ConvNets) have linked clinical flow cytometric data with patient responses56. Additionally, an ensemble method integrating FlowSOM and ConvNet has been utilized for haematological B-cell neoplasm datasets7. However, these methods often require large sample sizes (hundreds to thousands) and extensive image data for training ML models, in contrast to the smaller datasets typically derived from certain flow cytometric animal studies, which can number less than ten or up to tens of samples. Moreover, existing methods designed for multidimensional data incorporating various marker protein expressions may not suitably analyse Timer fluorescence data without introducing bias.
[0008] The present invention overcomes these limitations by providing a classification approach utilizing a novel image conversion technique which is capable of producing effective input images enabling accurate and precise classification of cell samples. In turn, this classification provides unique insights into the transcription mechanisms of key proteins. Technologically, the approach of the present invention advances beyond traditional strategies including pre-determined gating and unsupervised clustering in flow cytometry, enabling a more sophisticated ML-assisted, data-oriented analysis of flow cytometric data.
[0009] The work underlying the present invention was supported by the Medical Research Council [grant number MR / S000208 / 1]; and the Biotechnology and Biological Sciences Research Council [grant number BB / J013951 / 2],
[0010] Summary of the invention
[0011] The present inventors have developed an approach for generating an input image for a classifier for classifying a sample of cells as one of a plurality of cell classes. The approach advantageously uses a novel method for converting fluorescence measurements into a 2D input image, to enable the classification of cells using an image classifier. Use of the method provides unique insights into transcriptional dynamics of a cell population that are not achievable with existing methods.
[0012] Aspects of the invention are set out in the independent claims. Optional features of the invention are set out in the dependent claims.
[0013] In a first aspect, the present invention provides a method of generating an input image for a classifier for classifying a sample of cells as one of a plurality of cell classes. The method of the first aspect comprises obtaining fluorescence measurements from each cell of a sample of cells. The cells of the sample have been adapted to express a fluorescent protein having expression associated with that of a protein of interest, and the fluorescent protein fluoresces in a composite colour. The composite fluorescent colour is decomposable into a first colour and a second colour. The method further comprises, for each measurement, determining an intensity of fluorescence in the first colour and an intensity of fluorescence in the second colour. A total intensity of fluorescence in the first and second colours combined is recorded, as well as a mixing value indicative of the ratio of the respective intensities of fluorescence in the first colour and fluorescence in the second colour. The input image is generated as a 2D image of pixel values with a first dimension corresponding to the total intensity and a second dimension corresponding to the mixing value, each pixel corresponding to a respective first range of the total intensity and a respective second range of the mixing value and a value of each pixel being indicative of a count of the fluorescence measurements having a total intensity in the first range and a mixing value in the second range.
[0014] In a second aspect, the present invention provides a method of classifying a sample of cells as one of a plurality of cell classes. The method of the second aspect comprises generating an input image according to the first aspect of the invention and inputting the input image to a classifier to obtain a classification output. The classifier has been trained to classify input images obtained according to the first aspect of the invention as one of the plurality of cell classes. A cell class is determined for the input image using the classification output.
[0015] In a third aspect, the present invention provides a method of training a classifier to classify a sample of cells as one of a plurality of cell classes. The method of the third aspect comprises generating a data set comprising for each of a plurality of samples a respective input image and a respective class label. The input image has been generated according to the first aspect of the invention and the class label is indicative of a class of the plurality of cell classes. The classifier is trained to output a respective class label in response to an input of an input image using the data set.
[0016] In a fourth aspect, the present invention provides one or more computer-readable media comprising coded instructions that, when run on a computing device, implement a method or a classifier of any of the preceding aspects of the invention.
[0017] In a fifth aspect, the present invention provides a system comprising a memory and one or more processors. The one or more processors of the system are configured to perform the method of any of the first, second or third aspects of the invention. Detailed description
[0018] In a first aspect, the present invention relates to a method of generating an input image for a classifier for classifying a sample of cells as one of a plurality of cell classes, comprising obtaining fluorescence measurements from each cell of the sample of cells and converting the measurements to a 2D input image of pixel values suitable for input into the classifier via a number of steps. Each of the steps are performed sequentially as illustrated in Figure 1, each step taking the output of the previous step and using it as an input.
[0019] The method advantageously allows novel insights into transcriptional dynamics of a protein of interest. The plurality of cell classes may comprise cell classes each having different states of transcriptional activation for the protein of interest, such that classification of the cells allows determination of what state of transcriptional activation the cell is in. The plurality of cell classes may therefore comprise a wild-type cell class having unedited genes and a cell class having eliminated or reduced expression of a transcription factor acting on the protein of interest. The transcription factor may be CNS2, which is a transcriptional activator of Foxp3.
[0020] The first step, 101, is obtaining fluorescence measurements from each cell of a sample of cells, wherein the cells of the sample have been adapted to express a fluorescent protein having expression associated with that of a protein of interest, wherein the fluorescent protein fluoresces in a composite colour and the composite fluorescent colour is decomposable into a first colour and a second colour.
[0021] Obtaining fluorescence measurements may be done by retrieving or accessing information or data from a storage medium or a data record if the fluorescence measurements have been prepared previously. The data record can be any repository where data is stored, including but not limited to, a dataset, a database, or a server. Alternatively, obtaining fluorescence measurements may comprise measuring fluorescence produced by each cell of the sample of cells. The fluorescence may be measured using a flow cytometer, or any other suitable device for measuring cell fluorescence. The fluorescence measurements for each cell may comprise multiple fluorescence measurements for each cell, taken at different fluorescence wavelengths. The multiple fluorescence measurements are typically taken at a single time point.
[0022] The cells of the sample of cells are adapted to express a fluorescent protein having expression associated with that of a protein of interest. The expression of the fluorescent protein is associated with that of the protein of interest such that whenever the cell expresses the protein of interest, it also expresses the fluorescent protein. Accordingly, when expression of the protein of interest commences, so does expression of the fluorescent protein. Likewise, when expression of the protein of interest is terminated, so is expression of the fluorescent protein.
[0023] The cells of the sample of cells may be any suitable biological cell, such as T-cells, for example CD4+ T-cells, CD8+ T-cells, regulatory T-cells, effector T-cells, memory T-cells, natural killer T-cells, or gamma delta T-cells. The CD4+ T-cells can be Th1 T-cells, Th2 T- cells, Th17 T-cells, or T follicular helper cells. The regulatory T-cells can be native regulatory T-cells, effector regulatory T-cells, or T follicular regulatory cells. The memory T-cells can be central memory T-cells, effector memory-T cells, or tissue resident memory T-cells. The CD4+ or CD8+ T-cells can be naive T-cells. A naive T-cell is a T-cell that has not yet encountered and been activated by a specific antigen, and therefore has not differentiated into their final state yet. The cells may be B-cells, for example plasma cells, memory B-cells, pre-B cells, mature B-cells, follicular B-cells, marginal zone B-cells, transitional B-cells, or B1 cells. The cells may also be innate lymphoid cells or dendritic cells.
[0024] The cells of the sample of cells may be adapted such that the protein of interest is expressed in a construct comprising the protein of interest in addition to the fluorescent protein. Fluorescent proteins for use in the invention can be mutants of mCherry. Fluorescent proteins for use in the invention can be the result of combination of two fluorescent proteins with different expression dynamics into a single fusion protein. Fluorescent proteins for use in the invention include fastFT, dsRedE5TIMER and TandemFP38’9. TandemFP is a fusion protein consisting of DsRed2 and sfGFP. The construct is typically a genetic construct, or an engineered assembly of genetic material in sequence, where the gene corresponding to the protein of interest is combined with the gene expressing the fluorescent protein in sequence such that expression of both the protein of interest and the fluorescent protein are activated and driven together. This can be achieved by having a single promoter driving expression of both proteins. When a cellular event causes expression of the protein of interest to begin, expression of the fluorescent protein begins at the same time, and therefore fluorescence can be measured following the cellular event. The construct may be a bacterial artificial chromosome (BAC). The cells may be adapted by delivering the construct into the cell using any suitable method, for example injection of a BAC into the pronucleus of an animal subject embryo.
[0025] The fluorescent protein fluoresces in a composite colour. By this is meant the fluorescence produced by the fluorescent protein can be decomposed into at least a first colour and a second colour. The composite colour of fluorescence of the fluorescent protein changes over time. Fluorescent proteins whose colour changes over time are known in the art and include fastFT, dsRedE5TIMER and TandemFP. The change in colour of the fluorescence from a first colour to a second colour can be such that the first colour of fluorescence has a first half-life time, whilst the second colour of fluorescence has a second half-life time. The first half-life time indicates the amount of time it takes for half of the fluorescence of a plurality of fluorescent proteins to change from the first colour to the second colour. The second half-life time indicates the amount of time it takes for half of the plurality of fluorescent proteins to decay such that they no longer emit any fluorescence. Where the fluorescent protein is fastFT, the first half-life time can be 4 hours, and the second half-life time can be 122 hours. The fluorescent protein can be a mutant of mCherry, such as fastFT as described in Subach et al.3.
[0026] Where the first colour of fluorescence has a first half-life time, and the second colour of fluorescence has a second half-life time, the ratio of one colour of fluorescence compared to the other colour of fluorescence can be used as a measure of time elapsed since expression of the protein began. Where the fluorescent protein is a non-tandem protein, such as fastFT or dsRedE5TIMER, this is because each time a new fluorescent protein is transcribed, it will emit the first colour of fluorescence until a given time point, at which point the colour changes, and over time as more individual proteins reach the colour change time point the overall fluorescence emitted will shift from the first colour to the second colour. Where the fluorescent protein is a tandem protein, such as TandemFP, the change of overall fluorescence colour emitted is caused by a difference between the maturation kinetics between two fluorescent proteins, such as sfGFP and DsRed2, that means as one fluorescent protein degrades quicker than the other, the overall colour shifts towards that of the slower degrading protein. This allows linking of expression of the specific protein to time, via the fluorescence of the fluorescent protein. The protein of interest is typically a marker protein, for example Nr4a3, Foxp3, IL2, or PD1. By marker protein is meant proteins that serve as indicators or markers of specific biological processes, cellular states, or responses. They may be associated with particular signalling pathways, cellular functions, or disease states, and their expression levels or activity can provide insights into underlying cellular mechanisms.
[0027] The second step, 102, is determining for each measurement an intensity of fluorescence in the first colour and an intensity of fluorescence in the second colour and recording a total intensity of fluorescence in the first and second colours combined and a mixing value indicative of the ratio of the respective intensities of fluorescence in the first colour and fluorescence in the second colour.
[0028] Determining an intensity of fluorescence in the first colour and an intensity of fluorescence in the second colour may be done using a flow cytometer. Typically, the flow cytometer will detect the fluorescence emitted by the fluorescent protein using specific filters which direct emissions having wavelengths within certain ranges to the appropriate detectors in the flow cytometer. The flow cytometer is therefore able to determine an intensity of fluorescence in the first colour and an intensity of fluorescence in the second colour. The intensity values may be normalised by any suitable method. The total intensity of fluorescence in the first and second colours combined and the relative mixing value indicative of the ratio of the respective intensities of fluorescence in the first colour and fluorescence in the second colour may be recorded by plotting a 2D graph of the intensity of fluorescence in the first colour against the intensity of fluorescence in the second colour for each fluorescence measurement. The total intensity of fluorescence in the first and second colours combined may be found by drawing a straight line from the origin of the graph to the point plotted for a given fluorescence measurement, and finding the length of the straight line. The mixing value for each measurement may be found by calculating the angle between either the straight line and the x axis of the graph, or the straight line and the y axis of the graph. The angles may be found using trigonometric transformation, using either sine or cosine functions. The mixing value may simply be a ratio of the x and y components (the respective intensities in the first and second colours) or any other suitable value indicative of the ratio.
[0029] The total intensity of fluorescence in the first and second colours combined indicates the strength of transcription of the fluorescent protein, and therefore also the strength of transcription of the protein of interest. The mixing value indicative of the ratio of the respective intensities of fluorescence in the first colour and fluorescence in the second colour indicates time elapsed following initiation of transcription of the fluorescent protein, and therefore also since initiation of transcription of the protein of interest. Accordingly, when combined, the combined fluorescence intensity and the ratio of the intensities of the two colours provide a measure of the strength of transcription and the duration of transcription.
[0030] The third step, 103, is generating as the input image a 2D image of pixel values with a first dimension corresponding to the total intensity and a second dimension corresponding to the mixing value, each pixel corresponding to a respective first range of the total intensity and a respective second range of the mixing value and a value of each pixel being indicative of a count of the fluorescence measurements having a total intensity in the first range and a mixing value in the second range.
[0031] The input image has two dimensions, a first dimension corresponding to the total intensity and a second dimension corresponding to the mixing value. The image is divided into a grid of pixels. Each pixel corresponds to a respective first range of the total intensity and a respective second range of the mixing value. The fluorescence measurements are binned into a given pixel if their total intensity and mixing value fall within the first range and the second range of the pixel, correspondingly. Accordingly, each pixel comprises a count of fluorescence measurements falling within the defined ranges of said pixel. The length and width of each pixel may be equidistant. The pixels may all have the same dimensions. The input image may comprise a pixel grid comprising at least 10, 25, 50, 100, 250 500 or 1000 pixels in each dimension. The input image may comprise a 100 x 100 pixel grid. The input image may be a two-dimensional histogram with intensity bins along one dimension, mixing value along the other dimension and a corresponding count of cells in each intensity I mixing ratio bin.
[0032] In a second aspect, as illustrated in Figure 2 the present invention provides a method of classifying a sample of cells as one of a plurality of cell classes. The method uses an input image generated according to the first aspect of the invention. The input image is inputted into a classifier to obtain a classification output (201). A cell class for the input image is determined using the classification output (202).
[0033] Many suitable image classifiers will be known to the person skilled in the art and the present disclosure is not limited to any particular one. Suitable classifiers include artificial neural networks and in particular convolutional neural networks. Image transformers are another class of neural networks that could be used. For the purpose of illustration rather than limitation, examples of artificial neural networks and convolutional neural networks are discussed below.
[0034] An artificial neural network (ANN) is a type of classifier that arranges network units (or neurons) in layers, typically one or more hidden layers between an input layer and an output layer. Each layer is connected to its neighbouring layers. In fully connected networks, each network unit in one layer is connected to each unit in neighbouring layers. Each network unit processes its input by feeding a weighted sum of its inputs through an activation function, typically a non-linear function such as a rectified linear function or sigmoid, to generate an output that is fed to the units in the next layer. The weights of the weighted sum are typically the parameters that are being trained. An artificial neural network can be trained as a classifier by presenting inputs at the input layer and adapting the parameters of the network to achieve a desired output, for example increasing an output value of a unit corresponding to a correct classification for a given input. Adapting the parameters may be done using any suitable optimisation technique, typically gradient descent implemented using backpropagation.
[0035] A convolutional neural network (CNN) may comprise an input layer that is arranged as a multidimensional array, for example a 2D array of network units for a grayscale image, a 3D array or three layered 2D arrays for an RGB image, etc, and one or more convolutional layers that have the effect of convolving filters, typically of varying sizes and / or using different strides, with the input layer. The filter parameters are typically learned as network weights that are shared for each unit involved in a given filter. Typical network architectures also include an arrangement of one or more non-linear layers with the one or more convolutional layers, selected from, for example pooling layers or rectified linear layers (ReLLI layers), typically stacked in a deep arrangement of several layers. The output from these layers feeds into a classification layer, for example a fully connected, for example non-linear, layer or a stack of fully connected, for example non-linear, layers, or a pooling layer. The classification layer feeds into an output layer, with each unit in the output layer indicating a probability or likelihood score of a corresponding classification result. The CNN is trained so that for a given input it produces an output in which the correct output unit that corresponds to the correct classification result has a high value or, in other words, the training is designed to increase the output of the unit corresponding to the correct classification. Many examples of CNN architectures are well-known in the art and include, for example, googLeNet, Alexnet, densenet101 or VGG-16, all of which can be used to implement the present disclosure.
[0036] The classification output may comprise a classification score. The classification output may comprise a classification score for each class of the plurality of cell classes. The classification score may indicate the likelihood (or log likelihood or logit) that the sample of cells belongs to a class of the plurality of cell classes. The classification score may indicate the likelihood that the sample of cells belongs to the cell class determined for the input image using the classification output. Determining a cell class for the input image using the classification output may comprise identifying which class has the greatest classification score, and determining that the sample of cells is most likely to belong to that predicted class.
[0037] The second aspect of the invention may further comprise determining a subset of pixels in the image whose value increases a classification score for the predicted class and determining a subset of cells within any one of the first and second ranges of the subset of pixels as cells contributing to the classification result. The subset of pixels may comprise fluorescence measurements being important or relevant for determining the class of the sample of cells. Where the classifier is a CNN, determining the subset of pixels may comprise using class activation mapping, such as gradient-weighted class activation mapping. Gradient-weighted class activation mapping typically comprises backpropagating a gradient of the classification score of the determined class with respect to a plurality of activations of a convolutional layer of the CNN. Typically, the convolutional layer I the final convolutional layer before an output layer of the CNN. Gradient-weighted class activation mapping may be carried out according to the method provided in Selvaraju et al.10. The subset of fluorescence measurements may be taken from feature cells that characterise the transcription dynamics of the protein of interest.
[0038] The classifier of the invention has been trained to classify input images obtained according to the first aspect of the invention as one of the plurality of cell classes. The classifier may be trained using the method of the third aspect of the invention.
[0039] In a third aspect, as illustrated in Figure 3 the present invention provides a method of training a classifier to classify a sample of cells as one of a plurality of cell classes. The method comprises generating a data set comprising for each of a plurality of samples a respective input image and a respective class label. The input image has been generated according to the first aspect of the invention. The class label is assigned to the input image to indicate which class of the plurality of cell classes applies to the sample (302).
[0040] The classifier is trained to output a respective class label in response to an input of an input image using the data set (303). Training the classifier may comprise providing input images of the data set as inputs to the classifier; obtaining outputs of the classifier in response to the images; comparing the outputs of the classifier with the corresponding class label for each input image to compute an error measure indicating a mismatch between the outputs and corresponding class labels; and updating parameters of the classifier to reduce the error measure. Any suitable training method may be used to train the classifier, for example adjusting the parameters using gradient descent.
[0041] Training may proceed for a number of epochs as described above until a classification error has reached a satisfactory level, or until the classifier has converged, for example as judged by reducing changes in the error between epochs, or the classifier may be trained for a fixed number of epochs. A proportion of the data set may be saved for evaluation of the classifier as a test dataset to confirm successful training. Once training is complete, the classifier, for example set of architecture hyperparameters and adjusted parameters, is stored 304 for future use. The adjusted parameters may be the network weights for an ANN or CNN classifier.
[0042] Figure 4 illustrates a block diagram of one implementation of a computing device 400 within which a set of instructions, for causing the computing device to perform any one or more of the methodologies discussed herein, may be executed. The computing device may also be referred to as a system, such as that of the fifth aspect of the present invention. In implementations, the computing device may be connected (e.g., networked) to other machines in a Local Area Network (LAN), an intranet, an extranet, or the Internet. The computing device may operate in the capacity of a server or a client machine in a client-server network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The computing device may be a personal computer (PC), a tablet computer, a set-top box (STB), a Personal Digital Assistant (PDA), a cellular telephone, a web appliance, a server, a network router, switch or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by that machine. Further, while only a single computing device is illustrated, the term “computing device” shall also be taken to include any collection of machines (e.g., computers) that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein.
[0043] The example computing device 400 includes a processing device 402, a main memory 404 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM), etc.), a static memory 406 (e.g., flash memory, static random access memory (SRAM), etc.), and a secondary memory (e.g., a data storage device 418), which communicate with each other via a bus 430.
[0044] Processing device 402 represents one or more general-purpose processors such as a microprocessor, central processing unit, or the like. More particularly, the processing device 402 may be a complex instruction set computing (CISC) microprocessor, reduced instruction set computing (RISC) microprocessor, very long instruction word (VLIW) microprocessor, processor implementing other instruction sets, or processors implementing a combination of instruction sets. Processing device 402 may also be one or more special-purpose processing devices such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), network processor, or the like. Processing device 402 is configured to execute the processing logic (instructions 422) for performing the operations and steps discussed herein.
[0045] The computing device 400 may further include a network interface device 408. The computing device 400 also may include a video display unit 410 (e.g., a liquid crystal display (LCD) or a cathode ray tube (CRT)), an alphanumeric input device 412 (e.g., a keyboard or touchscreen), and a cursor control device 414 (e.g., a mouse or touchscreen).
[0046] The data storage device 418 may include one or more machine-readable storage media (or more specifically one or more non-transitory computer-readable storage media) 428 on which is stored one or more sets of instructions 422 embodying any one or more of the methodologies or functions described herein. The instructions 422 may also reside, completely or at least partially, within the main memory 404 and / or within the processing device 402 during execution thereof by the computer system 400, the main memory 404 and the processing device 402 also constituting computer-readable storage media. The various methods described above may be implemented by a computer program. The computer program may include computer code (e.g. instructions) 510 arranged to instruct a computer to perform the functions of one or more of the various methods described above. The computer program and / or the code 510 for performing such methods may be provided to an apparatus, such as a computer, on one or more computer readable media or, more generally, a computer program product 500)), depicted in Figure 5. The computer readable media may be transitory or non-transitory. The one or more computer readable media 500 could be, for example, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, or a propagation medium for data transmission, for example for downloading the code over the Internet. Alternatively, the one or more computer readable media could take the form of one or more physical computer readable media such as semiconductor or solid state memory, magnetic tape, a removable computer diskette, a random access memory (RAM), a read-only memory (ROM), a rigid magnetic disc, and an optical disk, such as a CD-ROM, CD-R / W or DVD. The instructions 510 may also reside, completely or at least partially, within the memory 413 and / or within the controller circuitry 411 during execution thereof by a computing system 410, the memory 413 and the controller circuitry 411 also constituting computer- readable storage media.
[0047] In an implementation, the modules, components and other features described herein can be implemented as discrete components or integrated in the functionality of hardware components such as ASICS, FPGAs, DSPs or similar devices.
[0048] A “hardware component” is a tangible (e.g., non-transitory) physical component (e.g., a set of one or more processors) capable of performing certain operations and may be configured or arranged in a certain physical manner. A hardware component may include dedicated circuitry or logic that is permanently configured to perform certain operations. A hardware component may comprise a special-purpose processor, such as an FPGA or an ASIC. A hardware component may also include programmable logic or circuitry that is temporarily configured by software to perform certain operations.
[0049] In addition, the modules and components can be implemented as firmware or functional circuitry within hardware devices. Further, the modules and components can be implemented in any combination of hardware devices and software components, or only in software (e.g., code stored or otherwise embodied in a machine-readable medium or in a transmission medium). Unless specifically stated otherwise, as apparent from the following discussion, it is appreciated that throughout the description, discussions utilizing terms such as "generating”, “determining”, “recording”, “classifying”, “training”, or the like, refer to the actions and processes of a computer system, or similar electronic computing device, that manipulates and transforms data represented as physical (electronic) quantities within the computer system's registers and memories into other data similarly represented as physical quantities within the computer system memories or registers or other such information storage, transmission or display devices.
[0050] It is to be understood that the above description is intended to be illustrative, and not restrictive. Many other implementations will be apparent to those of skill in the art upon reading and understanding the above description. Although the present disclosure has been described with reference to specific example implementations, it will be recognized that the disclosure is not limited to the implementations described but can be practiced with modification and alteration within the spirit and scope of the appended claims. Accordingly, the specification and drawings are to be regarded in an illustrative sense rather than a restrictive sense. The scope of the disclosure should, therefore, be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled
[0051] The present invention is now further described by way of reference to the following Example which is present for the purposes of illustration only. Accompanying Figures are provided, in which:
[0052] Figure 1 : Illustrates a workflow for generating an input image according to the first aspect of the present invention.
[0053] Figure 2: Illustrates a process for classifying a sample of cells according to the second aspect of the present invention.
[0054] Figure 3: Illustrates a process for training a classifier to classify a sample of cells according to the third aspect of the present invention.
[0055] Figure 4: Illustrates a computer system on which the methods of the present invention can be implemented. Figure 5: Illustrates a computer program product on which the methods of the present invention can be implemented.
[0056] Figure 6: Experimental design of the disclosed example of the present invention.
[0057] Figure 7: Overview of machine learning approaches in the disclosed example of the present invention.
[0058] Figure 8: Development of CRISPR-mediated CNS2 KO in Foxp3-Tocky Mice. The Conserved Non-coding Sequence 2 (CNS2) of the Foxp3 gene was deleted specifically in the Foxp3-Timer transgene in Foxp3-Tocky using a CRISPR approach.
[0059] (a) Display of aligned sequence reads from ChlP-seq experiments using anti-Runx1 antibody in Foxp3-negative CD4+ T cells (Tconv) and Foxp3-positive CD4+ T cells (Treg), as well as anti-Foxp3 antibody in Treg. The RefSeq annotation data shows exons and introns. The CNS2 sequence is highlighted.
[0060] (b) The working model for CNS2-mediated Foxp3 transcriptional regulation. CNS2 is bound by both Foxp3 and Runxl proteins, which cooperatively regulate Foxp3 transcription. Notably, this model illustrates an autoregulatory loop where the Foxp3 protein itself participates in regulating its own transcription through interactions at the CNS2 site
[0061] (c) The CRISPR mutagenesis strategy used to delete CNS2. In the Foxp3 Timer bacterial artificial chromosome (BAC) transgene within the genome of Foxp3-Tocky mice, CNS2 was specifically targeted and excised to produce the CNS2 KO Foxp3 Timer transgene, hereafter designated as CNS2 KO Foxp3-Tocky. Notably, the first coding exon of the Foxp3 gene is replaced with the Timer gene in Foxp3-Tocky mice.
[0062] (d) The CRISPR transgenic approach to establish CNS2 KO Foxp3 Tocky. As a new sub-strain of the parental Foxp3-Tocky colony, CNS2 KO Foxp3 Tocky mice were developed by electroporating fertilized eggs from Foxp3-Tocky mice with CRISPR components targeting CNS2. This method ensured comparable genetic background, ensuring a reliable comparison to the parental strain. Founder mice were then screened to establish the novel CNS2 KO Foxp3-Tocky colony.
[0063] Figure 9: Analysis of Foxp3 timer fluorescence data using conventional methods. Flow cytometric Timer fluorescence data were obtained from Foxp3-Tocky mice (WT) and CNS2 KO Foxp3-Tocky mice (CNS2 KO).
[0064] (a) Representative flow cytometric 2D plot showing Timer Blue and Timer Red fluorescence from WT and CNS2 KO samples.
[0065] (b) Mean Fluorescence Intensity (MFI) analysis of KO and WT samples for Timer Blue (left) and Timer Red (right). Mann-Whitney test was used.
[0066] (c) Representative Timer Angle and Intensity plots after normalisation and trigonometric transformation of Timer Blue and Timer Red fluorescence, showing WT and CNS2KO samples.
[0067] (d) The Timer locus approach, categorising Timer Angle into the five loci: New (0°); New-to-Persistent-transitioning (NPt, 0° - 30°); Persistent (30° - 60°); Persistent-to- Arrested-transitioning (PAt) (60° - 90°); and Arrested (90°).
[0068] (e) Analysis of cell distribution by the Timer locus approach. The plots show the percentage of cells within each Timer locus in the parent population, specifically in CD4+ T-cells from the spleen (left) and superficial lymph nodes (right). Statistical significance was assessed using the Mann-Whitney test, with p-values adjusted by the Benjamini & Hochberg method. Significance levels are indicated as follows: p < 0.05 (*) and p < 0.01 (**).
[0069] Figure 10: Validation of TockyCNN using thymus-spleen datasets. This figure shows the performance of TockyCNN using training and test datasets analysing thymic and splenic cells from WT Foxp3 Tocky mice.
[0070] (a) Model training with learning curve. The TockyCNN model is trained using a threefold cross-validation method. (b) Receiver Operator Curve (ROC) analysis for the test dataset.
[0071] Figure 11 : TockyCNN ConvNet modelling using image conversion and gradient analysis.
[0072] (a) Image conversion technique as data preprocessing. Flow cytometric Timer fluorescence data are normalised and transformed via trigonometric transformation. Subsequently, Timer Angle and Intensity data are transformed into 2D image data, with a resolution of 100 x 100 pixels by binning, retaining cell density data.
[0073] (b) Image-converted Timer data are displayed by dot plots (upper) and pseudocolour plots (lower), with left plots showing concatenated Foxp3-Tocky (WT) samples and right plots showing concatenated CNS2 KO Foxp3-Tocky (KO) samples.
[0074] (c) The structure of the ConvNet designed specifically for analysing image-converted Timer data.
[0075] (d) Model training with learning curve. The TockyCNN model is trained using a threefold cross-validation method, focusing on constructing a CNS2-dependent Timer profile.
[0076] (e) Receiver Operator Curve (ROC) analysis for two test datasets.
[0077] (f) Gradient-weighted Class Activation Mapping (Grad-CAM) analysis of the TockyCNN model response to WT and CNS2 KO samples. Heat colour is used to show the reactive pixels by yellow highlights in each group. The dot plots for WT or KO are overlayed to show area of overlap.
[0078] (g) Differential heatmap analysis displaying the WT-specific response areas identified by Grad-CAM, shown in red heatmap colours.
[0079] (h) Feature cell identification by the TockyCNN model. Feature cells are visualised in the original Timer Angle and Intensity space (left) and Timer Blue and Red fluorescence space (right). (i) Analysis of WT feature cells. Percentage of WT feature cells by in the WT and CNS2 KO groups form test dataset 1 are displayed in a violin plot. The Mann- Whitney test was used to assess statistical significance, P < 0.001 (****).
[0080] (j) Mean Fluorescence Intensity (MFI) analysis of indicated markers for the WT feature cluster, other Timer+ cells, and Timer-negative cells. The test dataset 1 was used.
[0081] The Mann-Whitney test with p-value adjustment was applied, indicating statistical significance levels at P < 0.01 (**), P < 0.005 (***), P < 0.001 (****).
[0082] Figure 12: Testing TockyCNN model using test dataset 2.
[0083] (a) Gradient-weighted Class Activation Mapping (Grad-CAM) analysis of the TockyCNN model response to WT and CNS2 KO samples. Heat colour is used to show the reactive pixels by yellow highlights in each group. The dot plots for WT or KO are overlayed to show area of overlap.
[0084] (b) Differential heatmap analysis displaying the WT-specific response areas identified by Grad-CAM, shown in red heatmap colours.
[0085] (c) Feature cell identification by the TockyCNN model. Feature cells are visualised in the original Timer Angle and Intensity space (left) and Timer Blue and Red fluorescence space (right).
[0086] (d) Analysis of WT feature cells. Percentage of WT feature cells by in the WT and CNS2 KO groups form test dataset 2 are displayed in a violin plot. The Mann-Whitney test was used to assess statistical significance, P < 0.001 (****).
[0087] (e) Mean Fluorescence Intensity (MFI) analysis of indicated markers for the WT feature cluster, other Timer+ cells, and Timer-negative cells. The test dataset 2 was used. The Mann-Whitney test with p-value adjustment was applied, indicating statistical significance levels at P < 0.01 (**), P < 0.005 (***), P < 0.001 (****).
[0088] Example Biologically, the present study focusses on the transcriptional regulation of Foxp3, a key transcription factor for regulatory T-cells (Tregs)11. The Foxp3-Tocky system has revealed the highly dynamic nature of Foxp3 expression during both inflammation and homeostasis, uncovering an autoregulatory loop where Foxp3 transcription is modulated by the Foxp3 protein itself12. A critical element in this regulatory framework is the interaction between Foxp3 and Runxl , which not only influences Foxp3 protein function during Treg differentiation and suppressive activity13but also affects Foxp3 transcription14. It has been shown that both Foxp3 and Runxl bind to a key enhancer in the Foxp3 gene, Conserved Non-Coding Sequence 2 (CNS2), which is conserved across mammals including humans and mice15. CNS2 plays key roles in the maintenance of Foxp3 expression through epigenetic modifications15'16. This study addresses the specific temporal phases of Foxp3 transcription that are regulated by CNS2. Consequently, it aims to elucidate the CNS2-dependent, temporally dynamic mechanisms of Foxp3 transcription using the Foxp3-Tocky system and a CRISPR- based approach targeting CNS2.
[0089] Integrative Experimental Design for Studying Transcriptional Dynamics In Vivo Using Molecular Biology and Machine Learning
[0090] In this study, a new integrative approach was established to investigate transcriptional dynamics of Foxp3 in vivo that are controlled by a specific enhancer (Figure 6). In Foxp3-Tocky, the BAC Foxp3-Timer transgene reports Foxp3 transcriptional dynamics by Fluorescent Timer protein, which irreversibly and spontaneously changes its emission spectrum from a blue fluorescence to a red fluorescence with the maturation half-life 4.1 hours3. The mature red-form Timer protein is stable and its decay rate is 122 hours12. The major aim of this study was to understand Foxp3 expression dynamics through analysing Timer profiles. Given that Foxp3 expression shows dynamic changes12, it is essential to understand the dynamics at the single cell level.
[0091] However, even with the use of the Foxp3-Tocky system, the static nature of flow cytometric data and the limited dynamic range of the Timer protein presented significant challenges17. Additionally, there is no general mathematical solution for converting Timer profiles back to transcription dynamics. To address these issues, this study developed a machine learning (ML) approach, which is improved as it does not require mathematical or statistical assumptions about the nature of Foxp3 transcriptional dynamics or Timer profiles.
[0092] By generating a CRISPR mutant strain of the Foxp3-Timer transgene that lacks a specific enhancer region, flow cytometric data from WT and KO Foxp3-Tocky was used to identify enhancer-dependent temporal Foxp3 expression. In this approach, independent training and testing datasets were obtained by flow cytometric analysis of samples with a sample size, n = 10 - 20. While this may not be considered large in typical machine learning contexts, it is substantial for immunological studies involving animal subjects. Additionally, this number is technically at the high end, considering the practical and logistical challenges inherent in such studies (Figure 6). Next, an ML classifier was trained to differentiate between WT and KO samples based solely on their flow cytometric Timer profiles. Importantly, the behaviour of the model was monitored using an appropriate method, allowing for the identification of feature cells that the ML model uses to discriminate between the two groups. The characterization of these feature cells uncovered aspects of the temporal regulation of Foxp3 expression via the enhancer, which cannot be revealed by other methods.
[0093] Overview of Novel ML Approach to Analyse Flow Cytometric Tocky Data
[0094] To implement the research framework as presented in Figure 6, a new ML toolkit was developed which was designed to analyse flow cytometric Timer profiles and identify feature cells that distinguish between group classes (Figure 7). Data preprocessing included normalization and a trigonometric transformation approach, which converts Timer blue and red fluorescence data into Timer Angle and Timer Intensity (norm), as previously described2. Timer Angle and Intensity data were converted into a two- dimensional image dataset through binning in two-dimensional space. This image conversion allowed for the application of ConvNet for group classification. Model behaviour was monitored by Grad-CAM, which highlights significant regions in input images that contribute to the network's predictions10. This ML approach provided independent insights into the properties of enhancer-dependent transcriptional dynamics as measured by Foxp3-Tocky.
[0095] A Novel Experimental Tool to Investigate the Roles of Conserved Non-Coding Sequence 2 (CNS2) in Regulating Temporal Dynamics of Foxp3 transcription To identify a biologically significant enhancer sequence and establish a prototypic approach to studying Foxp3 transcriptional dynamics, Chromatin Immunoprecipitation sequencing (ChlP-seq) data was analysed. The analysis demonstrated that both Foxp3 and Runxl proteins uniquely bound to the CNS2 region (also known as the T reg-specific demethylation region) of the Foxp3 gene (Figure 8a) as reported previously15. Importantly, the analysis of temporal dynamics of Foxp3 transcription using Foxp3-Tocky revealed that Foxp3 protein is required for sustaining Foxp3 transcription12and that the CNS2 region is actively demethylated at the moment when Foxp3 transcription is sustained and persistent in the CD4 single-positive thymocytes2. Therefore, it was hypothesised that CNS2 functions as a platform for critical transcription factors, including Foxp3 itself and Runxl , to dynamically regulate the Foxp3 transcriptional activities. Deleting the CNS2 sequence would elucidate the temporal phases of Foxp3 transcriptional regulation that are dependent on CNS2 (Figure 8b).
[0096] The deletion of a sequence within a bacterial artificial chromosome (BAC) transgene could be done in vitro, followed by the creation of a new mouse strain. However, this approach would be subject to between-founder variations, which are significant in BAC reporter approaches18. Moreover, modifying the endogenous Foxp3 sequence could make any output reporter measurement secondary to the modified dynamics of the Foxp3 protein. Therefore, the CNS2 sequence was deleted within the BAC Foxp3- Tocky transgene only, without disturbing the endogenous Foxp3 gene (Figure 8c).
[0097] To achieve this ambitious aim, a CRISPR transgenic approach was utilized (Figure 8d). Fertilized eggs were obtained from Foxp3-Tocky mice, and electroporation- mediated CRISPR was performed (Materials and Methods), followed by founder screening and the establishment of a CNS2 KO Foxp3-Tocky colony. To facilitate the screening process, a single-stranded oligodeoxynucleotide (ssODN) was used that included homology arms to promote homologous recombination and replace the CNS2 sequence with a short oligo sequence (Figure 8e). The specific deletion of the CNS2 sequence in the BAC Foxp3-Tocky transgene, but not in the endogenous locus, was further confirmed by PCR and sequencing. Lastly, a stable line of CNS2 KO Foxp3-Tocky mice was established, inheriting the mutated Foxp3 Timer BAC transgene in a Mendelian manner, thus further excluding the involvement of the endogenous Foxp3 gene on the X chromosome.
[0098] Analysis of CNS2 KO Foxp3-Tocky Using Established Methods
[0099] Upon establishing the CRISPR mutant strain, the roles of CNS2 were investigated using previously established methods and flow cytometric analysis was performed on CD4+T-cells from Foxp3-Tocky (WT) and CNS2 KO Foxp3-Tocky (KO) mice (Figure 9a). The overall expression patterns in Timer Blue (immature) and Timer Red (mature) fluorescence showed only modest changes between KO and WT upon initial inspection, which posed a challenge for analysis using conventional approaches including manual gating. Mean Fluorescence Intensity (MFI) analysis revealed no significant differences in Timer Blue fluorescence between WT and KO, while Timer Red fluorescence was significantly lower in the KO group, although the difference was moderate (p < 0.01 , Figure 9b). These findings suggest that overall Foxp3 transcriptional activities are modestly attenuated in T-cells from KO mice.
[0100] Next, a previously reported method involving trigonometric transformation for quantitative analysis of Timer fluorescence profiles was employed (Figure 7)2. After normalisation and transformation, the Timer Angle and Intensity plot showed moderate changes between the two genotypes (Figure 9c). To quantitative analyse Timer Angle dynamics, the categorisation method known as the Timer Locus approach was used (Figure 9d). This approach categorises Timer Angle into five loci: New (0°), Persistent (30° - 60°), Arrested (90°), and transitional regions including NPt (New-to-Persistent transitioning, 0° - 30°) and PAt (Persistent-to-Arrested transitioning, 60° - 90°). Notably, the percentages of cells in NPt, Persistent, and PAt were decreased in CNS2 KO Foxp3-Tocky, both in lymph nodes and the spleen (Figure 9e). These results indicate that active Foxp3 transcription, as determined by relative high Timer Blue fluorescence, is attenuated in cells that have accumulated Timer Red protein in CNS2 KO.
[0101] Collectively, the established methods did not fully elucidate the types of Timer profiles controlled by CNS2. The major limitations of the existing methods include that the methods rely on inherently one-dimensional analyses, which do not allow for the examination of two-dimensional Timer data in their original form or in the transformed Timer Angle and Intensity space.
[0102] TockyCNN: A ConvNet Approach Using Grad-CAM for Feature Cell Selection Based on Timer Fluorescence Data
[0103] A ConvNet-based ML approach for classification-oriented feature cell selection was developed, aiming to provide deeper insights into CNS2-mediated regulation of Foxp3 transcription. It was suggested that converting Timer fluorescence data into image data may open new opportunities for ML applications. Importantly, this image conversion allows the full potential of ConvNet to be harnessed. Preliminary analysis indicated that using 100 x 100 'pixels' provided an excellent performance for classifying thymus and spleen samples from WT Foxp3-Tocky mice (Figure 10). After conversion, dot plot and pseudocolour representations of the Timer data ensured that the overall characteristics were preserved (Figure 11a - 11b).
[0104] To prevent overfitting, a relatively small ConvNet model was intentionally constructed, comprising only two convolutional layers and two dense layers (Figure 11c). ROC analysis using two independent test datasets demonstrated satisfactory performance, with AUC values exceeding 0.9 (Figure 11d). To monitor the model’s behaviour, the last convolutional layer was examined using Grad-CAM10'19. This analysis showed that Grad-CAM effectively visualized the areas of importance for the classification task. Notably, the WT Grad-CAM aligned well with the distribution of WT cells, and similarly, the KO Grad-CAM with that of KO cells (Figure 11e). These findings affirm that the trained ConvNet model successfully captured the features critical for classification. Heatmap visualizations of the differences between the two Grad-CAM results highlighted an area approximately in the centre of the Angle and Intensity image space (Figure 11f). Reverse conversion of the feature 'pixels' for WT samples into cells in the original Timer Angle and Intensity space allowed for the visualization of feature cells as a cluster in a unique region (Materials and Methods, Figure 11g). Furthermore, these WT feature cells, as shown in Figure 11g, were significantly reduced in the KO group (Figure 11h), exhibiting high CD25 expression similar to other Timer cells and higher PD1 expression than other Timer non-feature cells. Timer Blue fluorescence was significantly higher in the WT feature cell cluster compared to other cells, whereas Timer Red fluorescence was lower than in non- feature cells. Similar results were obtained by applying the TockyCNN model to the other organ, superficial lymph nodes, from the dataset (Figure 12a - 12e).
[0105] Collectively, these results support the utility of the TockyCNN modelling approach for identifying the CNS2-dependent Foxp3 expression profile.
[0106] Discussion:
[0107] This study offered unique biological insights into the roles of CNS2 in Foxp3 regulation, underscoring its critical function in T-cell biology. The previously reported Tocky Locus approach, which noted a nuanced decrease of T cells in NPt - Persistent - PAt loci in Timer Locus dynamics (Figure 4e), contrasts with our current findings where a machine learning approach identified key feature cells essential for efficient data-oriented classification of genotypes. I ntriguingly , feature cells were diminished in T cells from CNS2 KO Foxp3-Tocky mice, indicating that CNS2 uniquely influences distinct phases of Foxp3 transcriptional dynamics. Specifically, the TockyCNN model identified cells between 50°- 80° in Timer Angle and these cells were uniquely PD-1 high (Figure 7j). These findings support the hypothesis that CNS2 may be redundant at the new phase of Foxp3 expression but contributes to the maintenance of Foxp3 expression at the moment when T-cells actively transcribe Foxp3 (Persistent locus), and further enhances the transcriptional activities towards a high end of Foxp3 Timer Intensity profile.
[0108] Furthermore, the differential expression observed in the CNS2 KO versus WT groups, particularly regarding CD25 and PD1 expression, revealed the broader role of CNS2 in regulating Foxp3 activity and T-cell function. Importantly, CD25 constitutes the Interleukin (I L)-2 receptor and its expression can be induced by either T cell activation through T cell receptor signalling and / or IL-2 signalling20. In addition, Foxp3 promotes the transcriptional activities of CD2513. In contrast, PD1 expression is induced by TOR signalling as a negative feedback mechanism for TOR signalling. Given these backgrounds, the findings in the present study support the hypothesis that CNS2 facilitates sustained Foxp3 expression in response to TCR signalling. I ntriguingly, previous reports indicated the requirement of CNS2 for the maintenance of Foxp3+ cells in aged mice, Foxp3 expression over cell divisions in vitro15, and the high levels of Foxp3 expression in activated T reg, or ‘effector Treg’16. Furthermore, previous work showed that, by analysing thymic CD4-single positive cells from Foxp3-Tocky, CpG islands within CNS2 (i.e TSDR) start to be demethylated in Persistent Foxp3 expressors but not in T cells in the New- NPt loci12. These two lines of studies support the possibility that the active usage of CNS2 at the moment of persistent Foxp3 transcription promotes the demethylation of the CNS2 CpG islands, which is considered a hallmark of ‘stable Foxp3 expression’21.
[0109] Materials and Methods:
[0110] CRISPR-mediated mutagenesis of Foxp3-Tocky
[0111] The Foxp3-Tocky mouse strain (BAC Tg (Poxp3AExon3 FastFT) was generated and reported previously2. Briefly, for generating the transgenic construct, a bacterial artificial chromosome (BAC) approach was used to replace the first coding exon of the Foxp3 gene with Fast-FT3using a knock-in knock-out approach. Subsequently, the BAC construct was injected into pronucleus at the two-cell stage of embryos using C57BL / 6J mice. All animal experiments were approved by the Animal Welfare and Ethical Review Body at Imperial College London and the Animal Experiment Committee at Kumamoto University.
[0112] The CRISPR-mediated mutagenesis of the Foxp3-Timer transgene in Foxp3-Tocky mice was performed using the previously reported CRISPR transgenic approach22 23. Briefly, fertilised eggs were obtained from Foxp3-Tocky and were subjected to electroporation to introduce the Cas9 protein (317-08441 , Nippon Gene, Toyama, Japan), tracrRNA (GE-002, FASMAC, Kanagawa, Japan), synthetic crRNAs, and a single-stranded oligodeoxynucleotide (ssODN). The CRISPR strategy was designed to replace the CNS2 sequence with a short insert sequence, AAGTTTAAAC. Accordingly, the following synthetic crRNAs were used: (1) 5’ crRNA targeting CCTGAGCTCCATTATGACAG (PAM: AGG) and (2) 3’ crRNA targeting AGTTCCACAAGTATTTAAGG (PAM: AGG). The ssODN designed for homology- directed repair and the insertion of the short nucleotide sequence AAGTTTAAAC / GTTTAAACTTC was 5’-
[0113] CTCTTTATGTTTGGTCAGAACTTATAAGAAATCTCCTCCTGTTTAAACTTCAGAGG ATTGGAAAACCCTCTACTGTCCTGATCTGGGGTC-3’ (Figure 8e). ChlP-seq Analysis
[0114] ChlP-seq data analysis was performed using the NCBI SRA dataset DRP00337624. Briefly, HISAT2 was used to align sequence reads to the mouse genome (mm10), followed by the data processing using samtools to produce sorted bam files. The sequence peaks were visualised by Integrative Genomics Viewer.
[0115] Timer Trigonometric Transformation and Normalisation
[0116] Data pre-processing for normalisation and trigonometric transformation was built on previously established methods2. Briefly, Timer fluorescence data were normalised across the immature blue and mature red fluorescence data, so that the two types of Timer fluorescence were treated equally. Subsequently, Timer fluorescence data were thresholded for removing autofluorescence, which occupies a large part of flow cytometric data. This was followed by trigonometric transformation of normalised timer blue and red fluorescence, converting Timer data into Timer-Angle and Timer-Intensity2. The R package flowCore25was used for importing fcs data.
[0117] TockyCNN: A Convolutional Neural Network Model for Flow Cytometric Timer Fluorescence Data
[0118] In this study, a unique ConvNet model, TockyCNN, was developed to analyse two- dimensional Timer fluorescence flow cytometric data (Figure 11c). By integrating an image conversion technique, TockyCNN model leverages image analysis techniques in ConvNets and thereby aims to improve the accuracy of classifying cellular states based on their transcriptional activity as reported by the Foxp3-Tocky system.
[0119] For data preprocessing, the blue-red normalisation and trigonometric transformation of Timer fluorescence data into Timer Angle and Intensity were employed. Subsequently, Timer Angle and Intensity data were converted into image data as follows. First, the global minimum and maximum values for the selected markers, corresponding to angle and intensity measurements, were computed. These values were employed to define uniform bin edges across all datasets, facilitating a standardized pixel representation. Specifically, the bin edges for angle and intensity were established using a sequence of 100 equidistant intervals spanning from the global minimum to the global maximum for each marker. Using these predefined bin edges, data were transformed into pixel-style representations, mapping the flow cytometric measurements into a 2D matrix. Each matrix element quantifies the count of data points within each bin, thereby converting the original flow cytometric data into a 100 x 100 pixel image where each pixel corresponds to a bin count. This method ensured a consistent and precise pixelated representation of the flow cytometric data for each sample.
[0120] To construct a ConvNet model tailored for Timer Angle and Intensity image data, TensorFlow and Keras were employed. The network features a series of convolutional layers enhanced with spatial attention mechanisms to focus processing on pertinent areas within the images. The input layer accepts batch image data of size 100 x 100 pixels with a single channel representing cell density. The initial convolutional layer, employing 16 filters of size 3 x 3 with 'same' padding, maintains dimensionality and is paired with a ReLU activation function to introduce non-linearity. The convolution is followed by a spatial attention block, comprising a 1 x 1 convolution with sigmoid activation to generate an attention map. This map is then multiplied element-wise with the input feature map to emphasise significant features. A subsequent max pooling layer reduces spatial dimensions, followed by a dropout layer at a rate of 0.2. This configuration, including a convolutional layer, an attention block, max pooling, and dropout, is repeated twice. The convolutional layers feed into a flatten layer that transitions the data for classification. This is followed by a fully connected dense layer with ReLU activation and a final dropout at the rate of 0.5 before reaching the softmax output layer, which classified the cells into two categories based on their transcriptional activity.
[0121] The model training utilized the Adam optimizer. To address potential class imbalances, class weights were calculated using the compute_class_weight function from scikit-learn26, ensuring equitable training by weighting the influence of each class according to its representation in the dataset. These computed weights were then integrated into the training process by passing them to Keras, enhancing the model's focus on underrepresented classes. Given the relatively small sample size, a three-fold cross-validation method was used during the training phase to ensure the robustness and generalizability of the model. For each fold, the model was compiled with the binary cross-entropy loss function and accuracy as the metric. Training was conducted over 12 epochs with a batch size of 8 and the learning rate 0.002, and the validation data for each fold was used to monitor model performance. The training histories were recorded for each fold to analyse model behaviour across different iterations.
[0122] Implementation of Grad-CAM in TockyCNN
[0123] Grad-CAM10was employed to identify critical feature cells within flow cytometric Timer fluorescence data. This method aided in distinguishing between WT Foxp3- Tocky and CNS2 KO Foxp3-Tocky samples using a pre-trained TockyCNN model. To implement Grad-CAM, a new gradient model was created using TensorFlow's Keras API. This model shared the same inputs as the original but outputs both the activations from the last convolutional layer and the model's final output. The tf.GradientTape function records operations for automatic differentiation and computes gradients of the model's output with respect to the last convolutional layer. These gradients indicate the importance of each neuron in the convolutional layer for class prediction. By averaging the gradients across the width, height, and depth of the feature map, a single scalar value per feature map channel is obtained. Subsequently, a weighted combination of the feature maps with the pooled gradients is calculated, resulting in a 2D array. Each element in this array is the product of the feature map activation and the importance of that feature map, thereby creating a raw heatmap. A ReLU function is applied to the heatmap to retain only positive values, which are then normalized to a range between 0 and 1 by dividing by the maximum value in the heatmap. After generating individual heatmaps for each image in the KO and WT groups, the heatmaps were summed within each group to aggregate activation signals across multiple images. Each aggregated heatmap then underwent Gaussian smoothing to suppress noise and enhance the visibility of predominant activation patterns. Finally, a differential matrix was calculated by subtracting the KO group's Grad-CAM heatmap from the WT group's Grad-CAM heatmap. This differential heatmap effectively highlighted regions with distinct activation between the groups, identifying feature cells that contribute to the classification task performed by TockyCNN. References
[0124] 1 Barry, J.D., Dona, Erika, Gilmour, D. & Huber, W. TimerQuant: a modelling approach to tandem fluorescent timer design and data interpretation for measuring protein turnover in embryos. Development 143, 174-179 (2016).
[0125] 2 Bending, D. et al. A timer for analyzing temporally dynamic changes in transcription during differentiation in vivo. Journal of Cell Biology 217, 2931-2950 (2018).
[0126] 3 Subach, F.V. et al. Monomeric fluorescent timers that change color from blue to red report on cellular trafficking. Nat Chem Biol 5, 118-126 (2009).
[0127] 4 Duetz, C. et al. Computational flow cytometry as a diagnostic tool in suspected-myelodysplastic syndromes. Cytometry A 99, 814-824 (2021).
[0128] 5 Arvaniti, E. & Claassen, M. Sensitive detection of rare disease-associated cell subsets via representation learning. Nature Communications 8, 14825 (2017).
[0129] 6 Hu, Z., Tang, A., Singh, J., Bhattacharya, S. & Butte, AJ. A robust and interpretable end-to-end deep learning model for cytometry data. Proceedings of the National Academy of Sciences 117, 21373-21380 (2020).
[0130] 7 Dinalankara, W., Ng, D.P., Marchionni, L. & Simonson, P.D. Comparison of three machine learning algorithms for classification of B-cell neoplasms using clinical flow cytometry data. Cytometry Part B: Clinical Cytometry n / a.
[0131] 8 Yau, B. et al. A fluorescent timer reporter enables sorting of insulin secretory granules by age. J. Biol. Chem. 295 (27), 8901-8911 (2020).
[0132] 9 Khmelinskii, A. et al. Tandem fluorescent protein timers for in vivo analysis of protein dynamics.
[0133] Nature Biotechnology 30, 708-714 (2012). 10 Selvaraju et al. Grad-CAM: Visual Explanations from Depp Networks via Gradient-based Localization. International journal of computer vision 128, 336-359 (2020).
[0134] 11 Ono, M. Control of regulatory T-cell differentiation and function by T-cell receptor signalling and Foxp3 transcription factor complexes. Immunology 160, 24-37 (2020).
[0135] 12 Bending, D. et al. A temporally dynamic Foxp3 autoregulatory transcriptional circuit controls the effector Treg programme. The EMBO journal 37, e99013 (2018).
[0136] 13 Ono, M. et al. Foxp3 controls regulatory T-cell function by interacting with AMLl / Runxl. Nature 446, 685-689 (2007).
[0137] 14 Kitoh, A. et al. Indispensable role of the Runxl-Cbfp transcription complex for in vivo- suppressive function of FoxP3+ regulatory T cells. Immunity 31, 609-620 (2009).
[0138] 15 Zheng, Y. et al. Role of conserved non-coding DNA elements in the Foxp3 gene in regulatory T- cell fate. Nature 463, 808-812 (2010).
[0139] 16 Li, X., Liang, Y., LeBlanc, M., Benner, C. & Zheng, Y. Function of a Foxp3 cis-element in protecting regulatory T cell identity. Ce / / 158, 734-748 (2014).
[0140] 17 Ono, M. Unraveling T-cell dynamics using fluorescent timer: Insights from the Tocky system. Biophysics and Physicobiology advpub (2024).
[0141] 18 Van Keuren, M.L., Gavrilina, G.B., Filipiak, W.E., Zeidler, M.G. & Saunders, T.L. Generating transgenic mice from bacterial artificial chromosomes: transgenesis efficiency, integration and expression outcomes. Transgenic Res 18, 769-785 (2009).
[0142] 19 Chollet, F. Deep Learning with Python. Manning Publications Co., 2017.
[0143] 20 Ono, M. & Satou, Y. Spectrum of Treg and Self-Reactive T cells: Single Cell Perspectives from Old Friend HTLV-1. Discovery Immunology (2024). 21 Cameron, J., Martino, P., Nguyen, L. & Li, X. Cutting Edge: CRISPR-Based Transcriptional Regulators Reveal Transcription-Dependent Establishment of Epigenetic Memory of Foxp3 in Regulatory T Cells. The Journal of Immunology 205, 2953-2958 (2020). 22 Suzuki, H., Kinoshita, G., Tsunoi, T., Noju, K. & Araki, K. Mouse Hair Significantly Lightened
[0144] Through Replacement of the Cysteine Residue in the N-Terminal Domain of Mclr Using the CRISPR / Cas9 System. Journal of Heredity 111, 640-645 (2020).
[0145] 23 Takemoto, K. et al. Meiosis-Specific C19orf57 / 4930432K21Rik / BRMEl Modulates Localization of RAD51 and DMC1 to DSBs in Mouse Meiotic Recombination. Cell Reports 31 (2020).
[0146] 24 Kawakami, R. et al. Distinct Foxp3 enhancer elements coordinate development, maintenance, and function of regulatory T cells. Immunity 54, 947-961. e948 (2021). 25 Hahne, F. et al. flowCore: a Bioconductor package for high throughput flow cytometry. BMC
[0147] Bioinformatics 10, 106 (2009).
[0148] 26 Pedregosa, F. et al. Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research 12, 2825-2830 (2011).
Claims
CLAIMS1. A method of generating an input image for a classifier for classifying a sample of cells as one of a plurality of cell classes, comprising: obtaining fluorescence measurements from each cell of a sample of cells, wherein the cells of the sample have been adapted to express a fluorescent protein having expression associated with that of a protein of interest, wherein the fluorescent protein fluoresces in a composite colour, the composite fluorescent colour being decomposable into a first colour and a second colour; for each measurement determining an intensity of fluorescence in the first colour and an intensity of fluorescence in the second colour and recording a total intensity of fluorescence in the first and second colours combined and a mixing value indicative of the ratio of the respective intensities of fluorescence in the first colour and fluorescence in the second colour; and, generating as the input image a 2D image of pixel values with a first dimension corresponding to the total intensity and a second dimension corresponding to the mixing value, each pixel corresponding to a respective first range of the total intensity and a respective second range of the mixing value and a value of each pixel being indicative of a count of the fluorescence measurements having a total intensity in the first range and a mixing value in the second range.
2. A method of classifying a sample of cells as one of a plurality of cell classes, comprising: generating an input image according to claim 1; inputting the input image to a classifier to obtain a classification output, wherein the classifier has been trained to classify input images obtained according to claim 1 as one of the plurality of cell classes; and, determining a cell class for the input image using the classification output.
3. A method of training a classifier to classify a sample of cells as one of a plurality of cell classes, comprising: generating a data set comprising for each of a plurality of samples a respective input image and a respective class label, wherein the input image has been generated according to claim 1 and the class label is indicative of a class of the plurality of cell classes; and,training the classifier to output a respective class label in response to an input of an input image using the data set.
4. The method of claim 2, further comprising: determining a subset of pixels in the image whose value increases a classification score for the cell class determined for the input image; determining a subset of fluorescence measurements within any one of the first and second ranges of the subset of pixels as cells contributing to the classification result.
5. The method of claim 4, wherein the classifier is a convolutional neural network (CNN) having a plurality of convolutional layers and determining the subset of pixels comprises using class activation mapping.
6. The method of claim 5, wherein the class activation mapping is gradient-weighted class activation mapping.
7. The method of claim 6, wherein determining the subset of pixels comprises backpropagating a gradient of the classification score of the predicted class with respect to a plurality of activations of a convolutional layer.
8. The method of claim 7, wherein the convolutional layer is the final convolutional layer before an output layer of the CNN.
9. The method of any of claims 1 to 4, wherein the classifier is an artificial neural network.
10. The method of claim 9, wherein the artificial neural network is a convolutional neural network.
11. The method of any preceding claim, wherein the mixing value is a trigonometric function of the ratio of the respective intensities of fluorescence in the first colour and fluorescence in the second colour.
12. The method of any preceding claim, wherein obtaining fluorescence measurements comprises measuring fluorescence produced by each cell of the sample of cells.
13. The method of claim 12, wherein measuring fluorescence produced by each cell of the sample of cells is carried out using a flow cytometer.
14. The method of any of claims 1 to 11, wherein obtaining fluorescence measurements comprises accessing fluorescence measurements on a data record.
15. The method of any preceding claim, wherein the fluorescent protein having expression associated with that of a protein of interest is expressed in a construct comprising the protein of interest in addition to the fluorescent protein.
16. The method of claim 15, wherein the construct is a bacterial artificial chromosome (BAG).
17. The method of any preceding claim, wherein the plurality of cell classes comprises cell classes each having different states of transcriptional activation for the protein of interest.
18. The method of claim 17, wherein the plurality of cell classes comprises a wild-type cell class and a cell class having eliminated or reduced expression of a transcription factor acting on the protein of interest.
19. The method of claim 18, wherein the transcription factor is CNS2.
20. The method of any preceding claim, wherein the protein of interest is a marker protein.
21. The method of claim 20, wherein the marker protein is selected from Nr4a3, Foxp3, IL2, and PD1.
22. One or more computer-readable media comprising coded instructions that, when run on a computing device, implement a method or a classifier as claimed in any preceding claims.
23. A system comprising a memory and one or more processors, wherein the one or more processors are configured to perform the method of any one of claims 1-21.
Citation Information
Patent Citations
Systems and methods for high-throughput image-based screening
US11788123B2
Information processing apparatus and information processing system
US20230071901A1
Information processing device, information processing method, program, microscope system, and analysis system
US20230243839A1