A method for identifying cellular heterogeneity
By employing CNNs trained on super-resolution images from STORM, the method addresses the challenge of identifying cellular heterogeneity, offering precise cell classification and feature identification with enhanced interpretability.
Patent Information
- Application Number
- PCT/EP2025/055589
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-24
- Filing Date
- 2025-02-28
- Publication Date
- 2025-09-04
AI Technical Summary
Current methods struggle to accurately identify cellular heterogeneity using single-molecule localization microscopy due to the complexity of cellular features and the lack of transparency in deep learning models, making it difficult to extract biologically meaningful patterns from pixel data.
A method utilizing artificial neural networks, specifically convolutional neural networks (CNNs), trained on super-resolution images from techniques like STORM, to classify cells and identify features associated with cellular heterogeneity, incorporating interpretability techniques such as occlusion-based methods to understand model decisions.
The method effectively classifies cells and identifies subcellular features indicative of cellular heterogeneity, providing accurate insights into cell types, differentiation states, chromatin organization, and disease progression, while enhancing model transparency.
Smart Images

Figure EP2025055589_04092025_PF_FP_ABST
Abstract
Description
[0001] A METHOD FOR IDENTIFYING CELLULAR HETEROGENEITY
[0002] Field of the invention
[0003] This invention relates to a method of identifying cells and cellular features. In particular, the invention relates to a method of applying artificial neural networks to identify cells and cellular features associated with cellular heterogeneity.
[0004] Background of the invention
[0005] Cell specification and cell phenotype heterogeneity are key determinants of many biological functions, such as differentiation, reprogramming of somatic cells into pluripotent state, cell transformation during cancerogenesis and cell changes after viral infection.
[0006] Using single-molecule localization microscopy methods (SMLM), and specifically, stochastic optical reconstruction microscopy (STORM), some cellular features not recognisable with conventional microscopy, such as nanoscale arrangement of the chromatin fibres, and nucleosome group arrangements (clutches), and distribution of RNA Pol II as well as nascent RNA in the nucleus, may be recognised. These features can be the key determinants of the somatic or stem cell state.
[0007] Overall, identifying cellular phenotypic heterogeneity can provide key information about biological functions, and the chromatin structure of each cell can be a clear proxy of this heterogeneity. Single-molecule localisation microscopy (SMLM) allows changes in nuclear nanostructures to be both visualized and quantified. Current methods of analysing single-molecule spatial distribution, such as cluster algorithms, are powerful at extracting the nuclear location and local density of the localizations. However, it is unclear how the spatial distributions and densities of these molecules can be used for identifying cell states.
[0008] The field of artificial intelligence (Al) has seen significant advancements in recent years, owing to the availability of large datasets and increased computational power. One particularly impactful development within Al is the emergence of deep learning (DL), a subfield that exploits artificial neural networks (ANNs) to extract and learn complex patterns and rules from data. Among the various forms of ANNs, convolutional neural networks (CNNs) have become the de facto standard in the field of computer vision and has been successfully applied to a wide range of medical and healthcare imaging applications. DL models have long been considered “black boxes” due to a lack of transparency and interpretability of their results. However, recent advances in the development of algorithms capable of explaining and interpreting the model’s results and choices have made it possible to understand which features contribute to a model’s output. This has led to the discovery of key biological features. In practice, however, the complexity of cellular heterogeneity presents challenges in identifying subcellular / cellular features that indicate cellular heterogeneity from a biological image.
[0009] The present disclosure addresses the shortcomings in the field by accurately classifying cells images and identifying subcellular / cellular features that indicate cellular heterogeneity.
[0010] Summary of the Invention
[0011] The disclosure relates to methods of identifying cellular heterogeneity, comprising of acquiring a plurality of images of one or more cells, applying a trained artificial neural network to classify one or more cells of the plurality of images, and identifying a feature associated with cellular heterogeneity in one or more classified cells of the plurality of images. Also described are methods for providing the test set to train an artificial neural network. The disclosure also relates to a method that accurately classify images and identify subcellular I cellular features that indicate cellular heterogeneity.
[0012] In the first aspect, the present invention provides a method for identifying cellular heterogeneity within a population of cells. The method comprising acquiring a plurality of images of one or more cells, defining a test set comprising one or more of the acquired images, providing the test set to a trained artificial neural network, applying the trained artificial neural network to classify one or more cells in one or more images of the test set, and identifying a feature associated with cellular heterogeneity in one or more classified cells of the one or more images of the test set.
[0013] In embodiments, the invention provides a method wherein the artificial neural network is trained by defining a training set comprising of one or more of the acquired images, defining a validation set comprising one or more of the acquired images that are not within the training set and optionally defining a test set comprising one or more of the acquired images that are not within the training set or the validation set. The method also comprises initialising weights associated with a plurality of filters of the artificial neural network, applying the plurality of filters to the training set and defining the weight of each of the plurality of filters, applying the plurality of filters having the defined weights to the validation set and assigning a classification and predicted value of one or more cells in the images of the validation set, applying a cost function to determine a loss value, wherein the loss value is defined as a difference between the predicted value of the assigned classification and the true value for the one or more cells in the images of the validation set. The method further comprises updating the weight of each of the plurality of filters of the artificial neural network, reapplying the plurality of filters to the validation set and re-assigning a classification and new predicted value of the one or more cells in the images of the validation set, re-applying the cost function to obtain a reduced loss value defined as a difference between the new predicted value and the true value for the one or more cells in the images of the validation set. In another embodiment, the reduced loss value reaches a minimised loss value, wherein the minimised loss value is obtained from re-applying the cost function.
[0014] In another aspect, there is provided a method for training an artificial neural network for identifying cellular heterogeneity within a population of cells, the method comprising: acquiring a plurality of images of one or more cells, and defining a training set comprising of one or more of the acquired images, defining a validation set comprising one or more of the acquired images that are not within the training set and optionally defining a test set comprising one or more of the acquired images that are not within the training set or the validation set. The method further comprises providing the training set to an artificial neural network to identify cellular heterogeneity in one or more cells of the training set, optionally applying the trained artificial neural network to classify one or more cells of the images of the test set, and optionally identifying the feature associated with cellular heterogeneity in one or more classified cells of the images of the test set.
[0015] In another aspect, the invention provides a method for identifying cellular heterogeneity within a population of cells, the method comprising acquiring a plurality of images of one or more cells, applying a trained artificial neural network to classify one or more cells of the plurality of images and identifying a feature associated with cellular heterogeneity in one or more classified cells of the plurality of images.
[0016] In an embodiment, the artificial neural network is trained according to the methods of the present disclosure.
[0017] In embodiments of any aspect, identifying cellular heterogeneity involves identifying one or more different cell types, differentiation states of one or more cells, degrees of chromatin organisation, subcellular I cellular features, one or more features responsive to drug treatment, one or more features indictive of an extent of infection or disease, and I or one or more features associated with cancer cell pluripotency.
[0018] In embodiments of any aspect, the acquired images are super resolution images or non-super resolution images.
[0019] In an embodiment, the super resolution images are acquired by Single-molecule localization microscopy (SMLM), Stochastic Optical Reconstruction Microscopy (STORM), Photoactivated Localization Microscopy (PALM), DNA points accumulation for imaging in nanoscale topography (DNA-PAINT) or a Stimulated emission depletion (STED) microscopy.
[0020] Embodiments of any aspect may further comprise generating a further plurality of images of one or more cells based on one or more of the acquired images. In some embodiments, the simulated images are generated from one or more randomly selected image of the plurality of images of step (a), and wherein a generative Al approach is used to generate the simulated images.
[0021] In embodiments of any aspect, the feature or features are indicative of one or more cell types and / or pathologies, one or more differentiation states of a cell, chromatin organisation in a cell, transcriptional activity, subcellular I cellular characteristics of a cell, a type of cancer and / or cancer progression of a cell, pluripotency of a cell, an extent of disease progression of a cell, response to drug treatment of a cell, and I or the viral infection level of a cell.
[0022] In embodiments of any aspect, the feature is a biological feature, an endogenous feature of a cell, or an exogenous feature of a cell.
[0023] In embodiments of any aspect, one or more molecule of interest in the population of cells are labelled prior to image acquisition, wherein the labelling involves immunolabelling, click chemistry, DAPI staining and / or binding of one or a plurality of fluorophores to the one or more molecules of interest.
[0024] In embodiments of any aspect, identifying a feature associated with cellular heterogeneity in one or more classified cells further comprises identifying an area of one or more images used for classification by applying a primary attribution algorithm and determining a feature of interest from the identified area of the one or more images.
[0025] In another aspect, there is provided a system for identifying cellular heterogeneity within a population of cells, comprising a database library that includes a training set comprising of one or more of the acquired images, a validation set comprising one or more of the acquired images that are not within the training set and a test set comprising one or more of the acquired images that are not within the training set or the validation set. The system further comprises a processor in communication with the database, the processor being adapted to receive and provide the test set to a trained artificial neural network, apply the trained artificial neural network to classify one or more cells in one or more images of the test set, and identify a feature associated with cellular heterogeneity in one or more classified cells of the one or more images of the test set.
[0026] It will be appreciated that any features of one aspect or embodiment of the invention may be combined with any combination of features in any other aspect or embodiment of the invention, unless otherwise stated, and such combinations are envisaged and are intended to be directly and unambiguously disclosed herein, and to fall within the scope of the present invention.
[0027] Brief Description of the Drawings Figure 1 AINU trained with Pol II and H3 images correctly identifies somatic cells and iPSCs. (a) Representative full-size, 10x rendered images of hiPSCs and somatic cells from dual-colour STORM localizations of Pol II (green) and H3 (red). White squares include representative patches reported in Figure 6a. (b-e) AINU trained with images rendered from dual-colour localizations of Pol II and H3 was challenged on a test set of 71 previously unseen images representing 20% of the whole data set. Normalized confusion matrix (with indication of positively and negatively predicted images in parentheses) (b) shows the performance of the model for each class with the main diagonal showing the accuracy for each class. The receiver operating characteristic (ROC) curve (c) shows the performance of the model at all classification thresholds and the value of the area under the ROC curve (AUC). (d,e) Precision and recall plots for the somatic (d) and hiPSC (e) classes reporting the overall average precision (AP)
[0028] Figure 2 AINU trained with Pol II images correctly identifies somatic cells and iPSCs. (a) Representative full-size super-resolution image of a hiPSC cell rendered at 10x magnification using Pol II channel of dual-color Pol II and H3 STORM localizations (see methods). Colors correspond to the normalized values of kernel density estimate for each pixel, (b-e) AINU trained with images of Pol II rendered from both dual-color and single-color Pol II localizations in hiPSCs and somatic cells was challenged on a test set of 146 previously unseen images representing 20% of the whole data set. Normalized confusion matrix (with indication of positively and negatively predicted images in parentheses) (b) shows the performance of the model for each class with the main diagonal showing the accuracy for each class. The receiver operating characteristic (ROC) curve (c) shows the performance of the model at all classification thresholds and the value of the area under the ROC curve (AUC). (d, e) Precision and recall plots for the somatic cell (d) and hiPSC (e) classes, showing the overall average precision (AP). (f) Same as (a) using Pol II single-color localizations, (g-j) Same as (b-e) for AINU trained with images from single-color Pol II localizations and tested on 74 previously unseen images (20% of the whole dataset), (k) Representative full-size super-resolution image of HeLa cell rendered at 10x magnification using dual-color localizations of H3 (red) and Pol II (green) (see Methods), (l-o) Same as (b-e) for AINU trained with dual-color localizations of H3 and Pol II challenged by substituting the hiPSC test images with 33 HeLa images.
[0029] Figure 3 AINU trained with DNA images correctly identifies different cell states. Representative full-size super-resolution images of somatic cells (a) and hiPSCs (b) rendered at 10x magnification using DNA localizations (see Methods). Colours correspond to the values of kernel density estimate for each pixel, (c-d) AINU trained with images rendered from singlecolour DNA localizations in hiPSCs and somatic cells was challenged on a test set of 36 previously unseen images that represented 20% of the whole dataset. Normalized confusion matrix (with indication of positively and negatively predicted images in parentheses) (c) shows the performance of the model for each class with the main diagonal showing the accuracy for each class. The ROC curve (d) shows the performance of the model at all classification thresholds and the value of the AUC. (e-g) Representative full-size super-resolution images of Mock-infected (e), 1 hour post-infection (hpi) (f) and 3 hpi (g) cells rendered using DNA localizations (see Methods). Colours correspond to the values of kernel density estimate for each pixel. (h,i) Normalized confusion matrix (with indication of positively and negatively predicted images in parentheses) (h) shows the performance of the model for each of the 3 classes with the main diagonal showing the accuracy for each class. The ROC curves (i) show the performance of the model at all classification thresholds for each class versus the other two as well as the corresponding value of the AUC. The micro- and macro-average ROC curves of the three classes and their corresponding AUC are also plotted.
[0030] Figure 4 Pol II localizations in the nucleoli are essential for hiPSC identification, (a) Positive and negative attribution scores for representative hiPSCs and somatic cells using the occlusion-based method (64x64 px size with stride of 32 px) for AINU trained with single-color Pol II images. The green regions correspond to areas whose occlusion decreases the probability of predicting the correct class, and that are therefore considered positive for the prediction, while red regions correspond to “distracting” areas whose occlusion increases the probability of predicting the correct class. (b,c) Images showing the class activation maps (CAM) for the hiPSC image in (a) and its overlay to the original image (c). (d) Same as (a) for AINU trained with dual-color images of Pol II and H3 in hiPSCs and somatic cells. (e,f) AINU trained with dual-color Pol II and H3 images was challenged on a test set of 71 previously unseen images having the nucleoli occluded by filling them with the background color. Normalized confusion matrix (with numbers of positively and negatively predicted images in parentheses), (e) shows the performance of the model for each class; the main diagonal reports the accuracy for each class. The ROC curve (f) shows the performance of the model at all classification thresholds and the value of the AUC. (g,h) Precision and recall plots for the somatic cell (g) and hiPSC (h) classes, reporting the overall average precision (AP). (i,j) Boxplots showing a significant difference (two-sided Mann-Whitney U test) between the median localizations per cluster (i) and median clusters area (j) in somatic (n = 22) and hiPSCs (n = 21) nucleoli, (k) Barplot showing significant difference (two-sided unpaired / -test, P-value = 0.0028) of (ss)RT-qPCR level for asincRNA transcribed by Pol II from rDNA IGS 28, between IMR90 cells and IMR90-derived hiPSCs (error bars represent standard deviation). Mean and SD values are shown for n = 3 independent experiments.
[0031] Figure 5 Identification of the best CNN model architecture and AINU training and validation accuracies using different datasets, (a) Diagram listing all the different datasets used throught the paper: (i,ii) dual-color images from 349 full-size nuclei of hiPSC and somatic images: 278 for train / validation and 71 for testing; (iii,iv) dual-color images from patches of hiPSC and somatic cells: out of 349, 279 nuclei for train / validation, and 70 nuclei for testing; (v) full-size nuclei images of hiPSC and somatic cells simulated using GAN with the 290 images of dataset i: 10,000 images, 8,000 for train, and 2,000 for validation; (vi.vii) 733 full- size real nuclei images rendered from both 349 H3 + RNA Polll channel of dual-color localizations and 384 Pol II single color localizations: 586 nuclei for train / validation, and 147 nuclei for testing; (viii.ix) single-color localizations from 384 full size nuclei of hiPSC and somatic images: 307 nuclei for train / validation, and 77 nuclei for testing; (x) single-color localizations from 33 full-size nuclei images of HeLa cells; (xi) dual-color localizations from 30 full-size nuclei images of HeLa cells; (xii, xiii) DNA localizations from 185 full size nuclei images of hiPSC and somatic: 148 nuclei for train / validation, and 37 nuclei for testing; (xiv, xv) DNA localization from 264 A594 cells infected by HSV-1 : 211 nuclei for train / validation, and 53 nuclei for testing, (b) Flowchart of the 5-fold procedure to select the best performing model architecture using dual-color SR images of Pol II and H3. The process starts by setting and keeping fixed all the (hyper)parameters, and then performing 5-fold CV for 11 different CNN architectures. The best performing one, based on the average accuracy and loss over the 5- folds, is then selected, and the following (hyper)parameter is tested over different values using 5-fold CV. Finally, the best performing architecture and (hyper)parameters are retrained once by randomly splitting the whole dataset in training / validation with a withheld test set, which is used to challenge the trained model using previously unseen images. The green squares connected by arrows highlight the best model or (hyper)parameters resulting from running the 5-fold CV at each step.
[0032] Figure 6 AINU trained with patches and simulated dataset and challenged with real images (a) Representative patches of 256x256-pixel size corresponding to the white square of the image in Fig. 1a rendered at 30x magnification, (b) Image comparison between full-size 10x real images rendered from localization of Pol II and H3, and dual-color images simulated using GAN. (c) Representative dual-colour Pol II and H3 simulated images generated by a trained GAN model using the DiffAugment method, (d-g) AINU trained with simulated images of Pol II and H3 was challenged on a test set of 59 previously unseen real images that represented 20% of the whole dataset, (d) Confusion matrix reporting normalized and not normalized (in parentheses) values shows the performance of the model for each class with the main diagonals showing respectively the accuracy and the number of positively predicted test samples for each class. The ROC curve (e) shows the performance of the model at all classification thresholds and the value of the AUC. (f,g) Precision and recall plots for the somatic cell (f) and hiPSC (g) classes, reporting the overall average precision (AP). (h-k) Same as (d-g) for AINU trained with 256x256-pixel-wide patches of dual-color images of Pol II and H3. Figure 7 AINU challenged using datasets excluding urine epithelial cells and BJ fibroblasts, (a-d) AINU trained with images rendered from dual-color localizations of Pol II and H3 that excluded urine epithelial cells (both somatic and hiPSCs), which were used only as final unseen image test set. (a) Confusion matrix reporting normalized and not normalized (in parentheses) values shows the performance of the model for each class with the main diagonals showing respectively the accuracy and the number of positively predicted test samples for each class. The ROC curve (b) shows the performance of the model at all classification thresholds and the value of the AUC. (c,d) Precision and recall plots for the somatic cell (c) and hiPSC (d) classes, reporting the overall average precision (AP). (e-h) Same as (a-d) for AINU trained with images rendered from dual-color localizations of Pol II and H3 that excluded BJ fibroblasts (both somatic and hiPSCs), which were used only as final unseen image test set.
[0033] Figure 8 Interpretable Al. (a) Occlusion-based method for Interpretable Al. A window of selected size (e.g., 32x32px) is moved with a specific interval (stride) across the whole image tested by the model in the horizontal and vertical direction. At each occlusion step, the model performs its prediction and compares the probability for the predicted class with that of the non-occluded image. If the probability decreases, the occluded region is assigned with a positive attribution score, whereby the higher the score, the greater the decrease in probability. In contrast, if the probability for the predicted class increases, the occluded region is assigned a negative attribution score, whereby the higher the negative score, the greater the increase of probability, (b-d) Positive (green) and negative (red) attribution scores for one representative hiPSC applying the occlusion-based method, with occlusion region of different sizes and strides as indicated in (a), to AINU trained with single-color Pol II images of hiPSCs and somatic cells. Left, representative super-resolution image of a hiPSC rendered at 10x magnification using single-color Pol II localizations. Middle, overlay of the rendered image with corresponding positive and negative attribution scores. Right, corresponding positive and negative attributions scores only. The green regions correspond to areas whose occlusion decrease the probability of predicting the correct class, which are considered positive for the prediction, while red regions correspond to “distracting” areas whose occlusion increase the probability of predicting the correct class, (e) Same as (b) for AINU trained with single-color DNA images of hiPSCs and somatic cells, (f) Representative unmodified super-resolution image rendered at 10x magnification using dual-color Pol II and H3 localizations used to challenge AINU, (g) The same as in (f) with the nucleolus replaced with cloned areas from the corresponding image, to show the decrease of AINU performance for hiPSC identification, (h) Representative single-color Pol II (magenta) rendered image of a cell stained with NOLC1 (green) to highlight the nucleolus region. Figure 8. Diagrams representing (a) dataset xvi, (b) dataset xvii and (c) dataset xviii. Dataset xvi consists of RNA Pol II phSer5 super resolution images of 370 preleukemic cells and 151 EV (empty vector, controls). Dataset xvii consists of RNA Pol II phSer5 super resolution images of 120 preleukemic cells. Dataset xviii consists of DAPI stained nucleic acid confocal microscopy images of 606 preleukemic cells and 268 EV (empty vector, controls).
[0034] Figure 9. AINU trained with RNA Pol II images of preleukemic cells, (a) Confusion matrix of AINU trained with dataset xvi, showing performance of the model with the main diagonal demonstrating the accuracy for each class, (b) ROC curve of AINU trained with dataset xvi showing the performance of the model at all classification thresholds and the value of the area under the ROC curve (AUC). (c) Confusion matrix of AINU trained with dataset xvii. (d) ROC curve and AUC of AINU trained with dataset xvii. (e) Confusion matrix of AINU trained with dataset xviii. (f) ROC curve and AUC of AINU trained with dataset xviii.
[0035] Figure 10. Cluster analysis of fluorescent RNA Pol II molecules, (a) Normalized density of localization per area of nucleus of EV (control) and MLL-AF9 (preleukemic) conditions for dataset xvi. (b) Median Nearest Neighbour Distance (NND) between clusters for dataset xvi. Each point represents a cell. Darker black bar represents the mean. Mann-Wittney test, ***p < 0.001.
[0036] Figure 11. Representative images of EV (empty vector, control) (left) and preleukemic (right) cells. Each color corresponds to a cluster. The smaller box represents the zoom in of the selected area.
[0037] Figure 12. Interpretable Al / GradCAM Analysis of Preleukemic Condition Prediction.
[0038] The model predicts the sample as "preleukemic" with high confidence (0.99 probability). Left Panel: Original microscopy image. This is the raw data input into the model, showing the cellular morphology. Middle Panel: Activation maps from selected model layers, indicating regions (in brighter colors) deemed significant by the network. This highlights the internal focus areas during the prediction process, aiding in understanding which features the layers find most relevant. Right Panel: Heatmap overlay on the original image from specific layers of the neural network, emphasizing regions that significantly impact the model’s decision, with intense coloration highlighting critical areas for the preleukemic prediction.
[0039] Detailed Description of the Invention
[0040] All references cited herein are incorporated by reference in their entirety. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs (e.g. in cell culture, molecular genetics, nucleic acid chemistry and biochemistry). Unless otherwise indicated, the practice of the present invention employs conventional techniques of chemistry, molecular biology, microbiology, recombinant DNA technology, and chemical methods, which are within the capabilities of a person of ordinary skill in the art. Such techniques are also explained in the literature, for example, J. Sambrook, E. F. Fritsch, and T. Maniatis, 1989, Molecular Cloning: A Laboratory Manual, Second Edition, Books 1-3, Cold Spring Harbor Laboratory Press; Ausubel, F. M. et al. (1995 and periodic supplements; Current Protocols in Molecular Biology, ch. 9, 13, and 16, John Wiley & Sons, New York, N. Y.); B. Roe, J. Crabtree, and A. Kahn, 1996, DNA Isolation and Sequencing: Essential Techniques, John Wiley & Sons; J. M. Polak and James O’D. McGee, 1990, In Situ Hybridisation: Principles and Practice, Oxford University Press; M. J. Gait (Editor), 1984, Oligonucleotide Synthesis: A Practical Approach, IRL Press; and D. M. J. Lilley and J. E. Dahlberg, 1992, Methods of Enzymology: DNA Structure Part A: Synthesis and Physical Analysis of DNA Methods in Enzymology, Academic Press. Each of these general texts is herein incorporated by reference.
[0041] In order to assist with the understanding of the invention several terms are defined herein.
[0042] The terms “nucleic acid”, “nucleotide”, “polynucleotide”, and “oligonucleotide” are used interchangeably and refer to a deoxyribonucleotide (DNA) or ribonucleotide (RNA) polymer, in linear or circular conformation, and in either single- or double-stranded form. The shorthand “nt” denotes herein a “nucleotide”.
[0043] The term “sequence” in the context of nucleic acids refers to the specific order of nucleotide bases along the polymer chain. The sequence of adenine (A), thymine (T) (uracil (U) in RNA), cytosine (C), and guanine (G) carries the genetic information in the form of a code. A sequence in relation to an oligonucleotide can refer to one or more contiguous nucleotides or bases.
[0044] The term “cellular heterogeneity” in the context of this application refers to the inherent diversity or variation among individual cells within a population of cells. Heterogeneous features can be shown in a form of cell types, phenotypes, functions, or behaviours that result from a combination of genetic, epigenetic, and environmental factors.
[0045] The term “differentiation” means the process of developing specialized structures and functions in biology. During differentiation, cells in tissues can adapt a specialised shape, size, and function. The differentiation process starts with undifferentiated cells, often called stem cells or progenitor cells, which have the potential to develop into different cell types. As cells differentiate, they acquire specific characteristics that enable them to perform particular functions within the body. For example, in multicellular organisms, cells may differentiate into nerve cells, muscle cells, blood cells, or other specialized cell types. The term “chromatin”, as used herein, refers to the complex of genomic DNA and proteins that forms chromosomes within the nucleus of eukaryotic cells. In addition to the genomic DNA, chromatin includes histones, DNA-binding factors (DBFs), the basal transcription machinery and its nascent transcripts, replication and repair machineries that copy and maintain DNA, and many other molecules that interact with any of these components. Chromatin generally exists in two forms: heterochromatin, which is highly condensed and is typically not transcribed; and euchromatin, which is less condensed and can be transcribed. Within chromatin, the DNA molecule wraps around a core of eight histone proteins to create tight loops known as nucleosomes with 20 to 50 bp of linker DNA connecting adjacent nucleosomes in a ‘beads on a string’ arrangement. In the next level of local compaction, nucleosomes coil around each other in a stack, to form a 30nm spiral (known as a solenoid or chromatin fibre). Longer range interactions between proteins and DNA in distant locations then, in turn, cause the chromatin fibre to loops and folds around itself to eventually produce a chromosome.
[0046] The term “nucleosome” means a basic structural unit of DNA packaging in eukaryotic cells, comprising of a segment of DNA wound around a histone protein core.
[0047] The term “nucleosome clutches” or “clutches” means arranged groups of nucleosomes. The groups of nucleosomes may be assembled in heterogeneous groups of varying sizes.
[0048] The term “pluripotent stem cell” refers to a type of stem cell that has the capacity to differentiate into many different cell types within the body. Pluripotent stem cells are more versatile than other types of stem cells and have the potential to develop into cells of all three germ layers: endoderm, mesoderm, and ectoderm.
[0049] The term “human induced pluripotent stem cell” refers to a type of pluripotent stem cell that is artificially derived from somatic cells, through reprogramming. Reprogramming involves introducing specific transcription factors into the cells, which causes them to revert to a pluripotent state with characteristics similar to embryonic stem cells.
[0050] As used herein, the term “super resolution microscopy” means techniques in optical microscopy that enable image acquisition of higher resolutions than limited by the diffraction limit.
[0051] The term “fluorophore” means a fluorescent compound that can re-emit light upon light excitation. Fluorophores typically contain several combined aromatic groups, or planar or cyclic molecules with several IT bonds.
[0052] The term “localization” in microscopy refers to the identified positions of individual fluorescent molecules or particles within an image. The term “single localization microscopy methods (SMLM)” means a form of super-resolution microscopy, wherein individual fluorescent molecules are imaged in isolation due to being very sparsely distributed, allowing for their localization to high precision and accuracy. Subsequent imaging of many such molecules over multiple frames allows for the construction of images with higher resolution than possible with conventional optical microscopy. Different forms of SMLM rely on distinct methods to achieve sufficient sparsity / isolation within each frame, while allowing sampling of the whole field of view during the entire acquisition sequence.
[0053] The term “stochastic optical reconstruction microscopy (STORM)” refers to a form of superresolution microscopy, wherein a subset of fluorophores within the sample is stochastically activated by a light source. This partial activation can be achieved by using low-intensity light to excite only a small fraction of the fluorophores at any given time. In STORM, the process of partial activation is repeated with different subsets of fluorophores being activated and imaged in each cycle. Reconstruction step then comprises of compiling the acquired positional information from multiple activation cycles to build a super resolution image.
[0054] “Machine learning”, as used herein, means use of computer systems that are able to learn and adapt without following explicit instructions, by using algorithms and statistical models to analyse and draw inferences from patterns in data.
[0055] A “machine learning model” refers herein to a computational system that is designed to make predictions or decisions based on data.
[0056] A “training set” is a subset of data used to train a machine learning model.
[0057] “(Cellular I subcellular) features” refer to the characteristics of cells that can be observed or described, and which can be used to identify a type of or condition of a cell. Such a feature can relate to parts of a cell or structures within a cell.
[0058] The term “artificial neural network” refers to Machine learning (ML) is a field of study in artificial intelligence concerned with the development and study of statistical algorithms that can effectively generalize and thus perform tasks without explicit instructions.
[0059] The term “artificial neural network” or “neural network” refers to a computational network or model built using principles of biological neuronal organisation, comprising of interconnected nodes organised in layers. The layers include an input layer, which receives input data, and hidden layers that are multiple intermediate layers that enable the network to learn complex representations, and an output layer produces the final output. A neural network can be trained using training set by adjusting the weights to optimise the performance of the network. The term “convolutional neural network (CNN)” refers to a specialised type of artificial neural network designed for processing and analysing visual data. CNNs are particularly effective in tasks related to computer vision, image and video recognition, and other visual pattern recognition applications. CNN uses convolutional layers, which apply convolution operations to input data, unlike other type of neural networks. These layers help capture local patterns and spatial hierarchies by learning features from small local receptive fields and aggregating them across the entire input. CNNs typically consist of multiple convolutional layers followed by pooling layers for down-sampling and fully connected layers for making predictions based on the learned features.
[0060] The term “preleukemic” refers to a stage or condition that precedes the development of leukemia.
[0061] Identifying cellular heterogeneity
[0062] Identifying cellular heterogeneity involves the process of distinguishing and characterizing the diverse subpopulations of cells within a larger cell population, tissue, or organism. Cellular heterogeneity can be manifested in various ways, including distinct cell cycle stages, cell types, cell morphologies, or any dimensions of cell measurements at a single cellular level. Cellular heterogeneity is not limited to phenotype differences between individual cells and can be extended to behavioural differences. Cellular heterogeneity underlies the complexity and variability present in biological systems at the cellular level and can be a key determinant of many biological functions, such as differentiation, reprogramming of somatic cells into pluripotent state, and cell transformation during cancerogenesis. Cellular heterogeneity can be induced by a variety of factors, including external agents such as viral infection or drug treatment. The complexity of cellular heterogeneity presents a challenge in its characterisation.
[0063] In some embodiments, identifying cellular heterogeneity involves determining the progression of a cancer, diagnosing a disease, determining the infection, identifying cell types, identifying differentiation states, identifying degrees of chromatin organisation, identifying subcellular / cellular features, identifying response to drug treatment, identifying degrees of infection, or identifying cancer cell pluripotency.
[0064] Cellular phenotypes in images are inherently unstructured, and in contrast to other probes or biological assays which deliver specific biological quantities, images are made of pixels. The main challenge in identifying cellular heterogeneity is to extract biologically meaningful patterns accurately and sensitively from pixel data.
[0065] Chromatin structure as an indicator of cellular heterogeneity
[0066] One possible way to identify cellular heterogeneity is to observe the difference in chromatin structure in cells. Chromatin structure, defined by the spatial organisation of chromatin comprising of DNAs and nucleosomes, can undergo changes, resulting in alteration of gene expression. Chromatin structure is often characterised by the degree of compaction of nucleosomes, and the accessibility of DNA to cellular machinery, both of which are influenced by modifications to DNA and histone proteins. Different chromatin structure of cells within a population can be one of the signs of cellular heterogeneity, as different chromatin structures indicate differences in gene expression patterns, which characterises cell type and cell fate decisions during development or in response to environmental signals.
[0067] One way to characterise chromatin structure may be to identify the arrangement of nucleosomes in groups called clutches. The density and number of nucleosomes per clutches have been characterised (Ricci et al., 2015). For example, nucleosome clutches contain a median number of approximately 6 nucleosomes in somatic cells. The size of the nucleosome clutches inversely correlates with the pluripotency grade of mouse ESCs (mESCs), as well as human induced pluripotent stem cells (hiPSCs). Further, upon in vitro differentiation of mESCs into mouse NPCs (mNPCs), clutch size increases (Ricci et al., 2015). The results suggest that nucleosome I clutch density correlates with the open or closed state of the chromatin fibre, as well as cell state I cell type.
[0068] For instance, transcriptionally active clutches are enriched for RNA Pol II while the nascent RNA nanodomains are spatially distributed between the clutches and RNA Pol II. The DNA itself can be imaged by STORM, which allows visualisation of the changes in the compaction of the clutch- associated DNA due to the different epigenetic states of somatic and stem cells. Alternatively, the inventors have found out that by staining the cells, e.g. with DAPI or other suitable stains, and using diffraction limited microscopy (e.g. confocal, widefield microscopy) it may be possible to visualise clutch-associated DNA and changes thereto. Al trained models (such as Al of the Nucleus - explained below) can identify the cells that are stained with DAPI, or any other suitable stain, such as, for example, Hoechst dye.
[0069] Furthermore, chromatin structure is known to indicate several cell states, including viral infection and cancer progression. For instance, infection by herpes simplex virus type 1 (HSV1) is known to profoundly alter the chromatin structure of host cells, resulting in host cell heterogeneity. An important hallmark of cancer is characterized by heterogeneous patterns of chromatin alterations indicative of cell heterogeneity. For instance, histone H1 loss drives lymphoma by disrupting the 3D chromatin structure. Cellular heterogeneity characterised by a nanoscale difference in degree of chromatin decompaction is correlated with different stages of carcinogenesis, in line with the notion that the structure of specific chromatin signatures reflects an altered transcriptional profile that is established during oncogenesis.
[0070] However, how the spatial distributions and densities of chromatin structure can be used to identify cellular heterogeneity in diverse population of cells was previously unknown. Further, the full landscape of cellular features associated with cellular phenotypic heterogeneity is also unknown.
[0071] The present disclosure provides, in various aspects and embodiments, a novel method of identifying cellular heterogeneity within a population of cells. According to the disclosure, the novel method further identifies a feature associated with cellular heterogeneity, wherein a feature can be a chromatin structure.
[0072] Super resolution microscopy and Single molecule localization microscopy (SMLM)
[0073] Super resolution microscopy refers to techniques in optical microscopy that enable image acquisition of higher resolutions than limited by the diffraction limit. Single localization microscopy methods (SMLM) is a subset of super-resolution microscopy, wherein individual fluorescent molecules are imaged in isolation, allowing for their localization to high precision and accuracy. SMLM methods usually employ conventional wide-field excitation and achieve super-resolution by separating of point spread functions (PSF) of individual fluorescent molecules. Separation of PSFs signal to noise ratio (SNR) by reducing the scatter of localisations resulting from imaging a fluorescent molecule many times. One way to achieve localisation of individual fluorescent molecules is to temporally separate fluorescent emissions of distinct molecules using photoswitching properties of fluorophores, where fluorescent molecules can switch between an active, excited state and inactive state. The photoswitching properties can be modulated, for instance, using laser irradiation or by controlling the chemical environment, among other methods. Single molecule localization microscopy (SMLM) comprises of a variety of techniques including Stochastic Optical Reconstruction Microscopy (STORM), Photoactivated Localization Microscopy (PALM), and Points Accumulation for Imaging in Nanoscale Topography (PAINT).
[0074] Stochastic optical reconstruction microscopy (STORM) involves stochastic activation of a subset of synthetic fluorophores within the sample activated by a light source. This partial activation can be achieved by using low-intensity light to excite only a small fraction of the fluorophores at any given time. In STORM, the process of partial activation is repeated with different subsets of fluorophores being activated and imaged in each cycle. Reconstruction step then comprises of compiling the acquired positional information from multiple activation cycles to build a super resolution image.
[0075] Fluorescence photoactivated localization microscopy (PALM) utilises fluorescent proteins that can be activated by UV illumination. Point accumulation in nanoscale topography (PAINT) is an SMLM technique that relies on fluorophores that switch between free diffusion and immobilization by binding to a target. The most prominent variant of PAINT is DNA-PAINT, where transient immobilization is achieved by hybridization of DNA strands.
[0076] SMLM techniques represent a major advantage over diffraction-limited imaging as they can determine the single cell level, nanoscale arrangements of chromatin fibers in cells with a resolution of ~20 nm. STORM has been used to quantify the density and number of nucleosomes per clutch and to characterize somatic and stem cells according to these characteristics (Ricci et al., 2015). STORM enables characterization of the nanoscale structure of chromatin fibers, for instance, changes in the compaction of the chromatin due to the difference in epigenetic states (Otterstrom, J. et al., 2019), and types of molecules enriched in transcriptionally active clutches (Castells-Garcia, A. et al., 2022).
[0077] SMLM techniques, combined with current methods of analyzing single-molecule spatial distribution such as cluster algorithms, enhance localization precision of the target molecule during data acquisition (Ouyang, W et al., 2018) and for semantic segmentation (Bilodeau, A. et al. 2022).
[0078] Use of SMLM techniques have provided unprecedented images of fine cellular structures, including protein distribution in brain (Dudok et al., 2015), nanostructures of a microtubule network (Hu et al., 2013), and dynamic processes within the cell membrane. Use of SMLM techniques has newly revealed detailed chromatin structure such as distribution of histone modifications on individual fibers (Franek et al., 2021), histone modifications on higher-order chromatin structures (Xu et al., 2021), and chromatin folding in disease state (Xu et al., 2020).
[0079] However, a major limitation of SMLM techniques is their limited throughput, resulting from the small field of view required for single-molecule localization and the long processing time required to obtain a high-quality super-resolution image of a single cell. The limitation poses a challenge in acquiring an amount of data that enables analytic methods to reliably identify cell heterogeneity or cell states. The limitation also makes it difficult to systematically identify features or subcellular / cellular structures indicative of the cellular heterogeneity.
[0080] The present disclosure reveals a novel method that allows use of a minimal amount of training data from SMLM imaging. In some embodiments, the number of images in each class of the training set is at least about 75, at least about 150, at least about 500, or at least about 800.
[0081] Use of machine learning and artificial neural network in identification of images
[0082] Machine learning is a field of artificial intelligence (Al) that involves the development of algorithms and models that enable computers to learn from data and make predictions or decisions without being explicitly programmed. The primary goal of machine learning is to create systems that can generalize patterns from data and adapt to new, unseen examples.
[0083] Image classification is one of the applications of machine learning, where the goal is to predict, automatically categorize or label images based on their content. One way that machine learning algorithms achieve image classification is by training a machine learning model on a dataset of labelled images, where the algorithm learns to identify patterns and features associated with different classes or categories. In image classification, an input to the machine learning algorithm may be an image including an object, and an output may be a class label or category.
[0084] A machine learning image processing system builds and trains multiple machine learning classifiers, including neural networks. The machine learning classifiers may accurately and automatically extract images and perform image processing to detect particular attributes of the extracted images and classify particular features of images relevant for a problem in hand.
[0085] Image processing using machine learning involves several steps, including image pre-processing, training the pre-processed images with the machine learning model, validation and fine tuning, and application. To enable these steps, the input dataset may be divided into the training set, validation set, and test set. The training set is the input data on which the machine learning model learns patterns and relationships by adjusting its parameters through the optimization process. The validation set is a separate subset of the data that is not used during the training but is used to tune hyperparameters and evaluate the model's performance during training. The test set is a new, independent subset of the data that is not used during training or hyperparameter tuning, providing an unbiased evaluation of the model's performance. Pre-processing step involves resizing, normalisation, and data augmentation of the input data. Then the input images are taken in a vector form into a machine learning model to be trained. A machine learning model, such as a classifier can perform this task by producing class scores that are a function of each pixel value. The class with the highest score is the predicted class, and then the algorithms of the classifier can be designed to alter the strength of the connections in the network to produce a desired signal flow. The strength, also known as a weighting, is adjusted during the training phase based on the differences between its predictions and the actual labels in the training data. Backpropagation and optimization algorithms are then used to iteratively update the model's parameters to minimize prediction errors. A separate validation dataset monitors the model's performance during training and prevents overfitting. Hyperparameters may be adjusted to fine-tune the model's performance.
[0086] A neural network is a model containing an interconnected group of processing elements or "neurons" that process information using a connectionist approach to computation. Neurons are organised into layers which typically include an input layer, one or more hidden layers, and an output layer. Neural networks are often used to model complex relationships between inputs and outputs or to find patterns within data. Typically, neural networks process data in a non-linear, distributed, parallel fashion. Often a neural network is an adaptive system that changes its structure during a learning phase. Functions are performed collectively and in parallel by the processing elements, rather than there being a clear delineation of subtasks to which various units are assigned. Generally, a neural network involves a network of simple processing elements that exhibit complex global behaviour determined by the connections between the processing elements and element parameters.
[0087] Neural networks consist of layers, and each layer plays a specific role in processing and transforming input data. The input layer receives the raw input data, and each neuron or node in the input layer represents a feature or attribute of the input data. Hidden layers are intermediate layers, in which each node is connected to every node in the previous layer, and these connections have associated weights that are adjusted during the training process. The output layer produces the final result or prediction of the neural network. The number of nodes in the output layer depends on the nature of the task. For example, in binary classification, there might be one node for each class, while in multiclass classification, there will be one node for each possible class. A neural network employed can be shallow, meaning the network has a small number of hidden layers, or deep, with multiple hidden layers. Shallow networks are simplerand computationally less intensive, making them suitable for certain tasks where the complexity of deep neural networks may not be necessary. A neural network used can have a skip connection, meaning bypassing one or more intermediate layers in deep neural networks.
[0088] Deep learning is a subfield of machine learning that focuses on the development and training of artificial neural networks, particularly deep neural networks. Deep neural networks are characterized by having multiple layers comprising of deep architectures. These deep architectures enable the models to learn hierarchical representations of data. In contrast to traditional machine learning which often requires manual feature engineering, deep learning models can automatically learn features from raw data. Convolutional Neural Networks (CNNs) are a type of deep learning model that have become the standard for computer vision and a wide range of medical and healthcare imaging applications. Convolutional Neural Networks include a series of layers such as convolutional layers, pooling layers and fully connected layers, which are designed to automatically and adaptively learn spatial hierarchies of features from input images. Convolutional layers apply multiple filters to the input image to extract image related characteristics such as edges, corners and textures. Pooling layers help to decrease computational complexity and control overfitting. Fully connected layers serve as a classifier that takes the extracted image related characteristics and maps them to the final output classes.
[0089] Overcoming obstacles in diffraction-limited imaging
[0090] Conventionally, deep learning models have been used to classify whole cell images using a training set of a significant number of diffraction-limited or non-SMLM microscopy images. A deep neural network trained on ±1500 conventional clinical histopathology images can classify the type of lung cancer in histopathology images, with performance similar to trained human pathologists (Coudray N et al., 2018). However, the use of diffraction-limited images prevents the access to details, finer textures, and smaller structures visible in the high-resolution output that can capture cellular heterogeneity. Further, during diffraction limited imaging, high-frequency components are lost during the original sampling process, and hence the amount of biological information in each pixel.
[0091] Advantageously, the present disclosure overcomes limitations of the diffraction limited imaging by providing, in various aspects and embodiments, super resolution images. In embodiments, the plurality of images can be super resolution images or non-super resolution images. In some embodiments, the deep learning model is used to overcome challenges in identifying cellular I subcellular features from diffraction limited imaging. The deep learning model may analyse the structures of DAPI labelled DNA or cellular components highlighted with fluorophore-assisted click chemistry. In embodiments, other suitable stains, such as, for example, Hoechst dye may be used for labelling of cellular components. Further, in some embodiments, the present disclosure involves additional analysis on the localizations of the individual fluorescent molecules I particles within an image, to assist understanding of particle arrangements at a scale finer than that which traditional light microscopy could resolve (Example 17).
[0092] The present disclosure provides, in various aspects and embodiments, a novel method for identifying cellular heterogeneity within a population of cells, comprising acquiring a plurality of images of one or more cells, defining a training set comprising of one or more of the acquired images, defining a validation set comprising one or more of the acquired images that are not within the training set and defining a test set comprising one or more of the acquired images that are not within the training set or the validation set, training an artificial neural network by initialising weights associated with a plurality of filters of the artificial neural network, applying the plurality of filters to the training set and defining the weight of each of the plurality of filters, applying the plurality of filters having the defined weights to the validation set and assigning a classification and predicted value of one or more cells in the images of the validation set, applying a cost function to determine a loss value, wherein the loss value is defined as a difference between the predicted value of the assigned classification and the true value for the one or more cells in the images of the validation set from step, updating the weight of each of the plurality of filters of the artificial neural network, re-applying the plurality of filters to the validation set and re-assigning a classification and new predicted value of the one or more cells in the images of the validation set, re-applying the cost function to obtain a reduced loss value defined as a difference between the new predicted value and the true value for the one or more cells in the images of the validation set, applying the trained artificial neural network to classify one or more cells in one or more images of the test set, and identifying a feature associated with cellular heterogeneity in one or more cells of the images of the test set.
[0093] Overcoming obstacles in super resolution imaging
[0094] Further, the present disclosure relates to the novel method that can overcome the size limitation of super resolution microscopy or SMLM dataset, by optionally generating thousands of images that augments the training set (Example 1). In some embodiments, further plurality of images in between 1000 and 20000, can be produced using generative Al approach including generative adversarial networks (GAN) or Variational autoencoders (VAE). Generative Adversarial Networks and Variational autoencoders are deep learning models that can generate synthetic training data. These generative Al models are capable of generating high-quality, realistic samples that resemble the distribution of real data, compensating for limited, noisy, or not fully representative true data distribution.
[0095] In some examples and embodiments, single-color Pol II SMLM imaging training dataset (Figure 5a, datasets vi, vii, viii and ix) is used to overcome technical challenges in simultaneous detection of multiple targets in the same sample in dual-color STORM imaging (Example 9). Use of single color imaging can save time and resources required to prevent potential cross-reactivity between different antibodies and spectral overlap of fluorophores. The model's performance closely paralleled that of the model trained with dual-color images.
[0096] Training deep learning models with SMLM data
[0097] The present disclosure, in aspects and embodiments, entails a novel strategy that combines deep learning and SMLM data of nuclear features for identifying cell heterogeneity. In some embodiments, deep learning involves the use of convolutional neural networks. In some embodiments, fluorescence intensity-based whole-imaged cells are used as training datasets for deep learning models. In some embodiments, the training dataset is indicative of nuclear structures, and specifically the nanoscale spatial distribution of core histones, Pol II and DNA structure. In embodiments, single molecule localised STORM of Ser 5-phosphorylated RNA Pol II, histone H3 or DNA, in human somatic cells, hiPSCs, human cells infected by HSV-1 , cancer cells, precancerous cells, or preleukemic cells are used as training datasets for deep learning models. In embodiments, SMLM data indicative epigenetic traits, such as the transcriptionally active or repressive H3Ac, H3K9me3 and H3K27me3 histone marks, combined with additional SMLM data with multiple nuclear features can be used to train CNN. The resulting trained model, named AINU (Al of the Nucleus), demonstrated near-perfect efficiency in differentiating between the various cell types, independent of the specific molecules imaged.
[0098] In Example 3, a CNN model was trained on chromatin clutch structure or on RNA Polymerase II (Pol II) nuclear distribution. The training dataset in this example are dual-color STORM images of nucleosome core histone H3 and Pol II, in human somatic cells and hiPSCs respectively. The super resolution localizations of the two targeted molecules were used to generate the full-size ground truth dataset, with images rendered at 10x magnification (Figure 1a, Figure 5a), partitioned into training, validation and withheld test sets.
[0099] In some aspects and embodiments, a plurality of SMLM images comprising of training dataset are applied to distinct CNN architectures, prior to the identifying the optimal machine learning architecture. In some aspects and embodiments, selecting the optimal artificial neural network involves training known artificial neural network model architectures using SMLM images and / or distribution of individual labels indicating positions of molecules such as H3 or Pol II. In some embodiments, distribution of individual labels is expressed in a graph network or a point cloud function. In other embodiments, the distribution of individual labels indicates qualities such as density or length or shape of the labelled molecule of interest.
[0100] In Example 4, on 11 distinct CNN architectures (Table 2), the dual-color full-size dataset of Example 3 including labelled histone H3 and Pol II, was split into training and validation sets in 80:20 ratio. The dataset is then trained employing a stratified 5-fold cross validation. In k-fold cross- validation, the dataset is divided into k subsets or folds, and the model is trained and evaluated k times using each subset as separate training sets. In some embodiments, the dataset division involves random sampling of the images. Further, the present disclosure demonstrates the robustness of AINU in evaluating unseen images from unseen cell types (Example 8). When AINU was trained with dataset i (Figure 5a) withholding all images of urine epithelial cells or BJ fibroblasts, AINU still performed well with a weighted accuracy and F1 score of 0.8 ± 0.11 when urine epithelial cells were withheld, and 0.79 ± 0.09 when BJ fibroblasts were withheld (Figures 7a-h, Table 8).
[0101] Evaluation metrics of deep learning architecture
[0102] During training of neural networks, validation accuracy and loss are common criteria used to evaluate the performance of a model. Validation accuracy is a metric that measures the proportion of correctly classified instances in the validation dataset. It provides an indication of how well the model generalizes to new, unseen data. A higher validation accuracy indicates better generalization performance. Validation loss quantifies the difference between the predicted values and the actual values in the validation set.
[0103] In some aspects and embodiments of the invention, selection of the deep learning architecture involves choosing a model with the highest average validation accuracy and smallest average loss. In embodiments, the average accuracy and loss are over the five folds per each architecture are calculated and compared. In Example 4, DenseNet-121 is shown to be the top performing CNN architecture in identifying somatic cells and hiPSCs, with an average validation accuracy of 92.26 and an average loss of 0.292 (Figure 5b, Table 2), and was used for subsequent analyses.
[0104] The present disclosure provides, in various aspects and embodiments, a method of optimizing hyperparameters for the selected deep learning architecture. In embodiments, hyperparameters of different values are tested on the selected model, wherein the hyperparameters include image size, image augmentation techniques, epochs, batch size, learning rate and a type of optimizer. Image augmentation techniques are critical to ensure the diversity and complexity of the augmented images used for training. For instance, by applying random rotations or flips to training images, the model becomes more resilient to variations in orientation and position. In embodiments, data augmentation techniques include flipping, random rotation, and color adjustments, were applied using the Albumentations library. Epoch defines the number of times a learning algorithm iteratively processes the entire training dataset during model training. The optimizer adjusts the parameters of a model in order to minimize the error or loss function during the training. The optimizer also determines how the model parameters should be updated in response to the computed gradients of the loss function with respect to those parameters. In embodiments, the loss function used may include: Cross Entropy Loss, Mean Squared Error Loss, and Binary Cross Entropy Loss. The gradients indicate the direction and magnitude of the steepest increase of the loss, and the optimizer uses this information to iteratively update the model parameter. In embodiments, types of optimizers comprise of: stochastic gradient descent, ADAM, ADAMW, Adadelta, and RMSprop. Learning rate, in the context of gradient-based optimization algorithms, determines the amount by which the model parameters are updated. Selecting an optimal learning rate is important as a high learning rate can cause the optimization algorithm to overshoot the minimum, potentially leading to divergence or oscillation, and low learning rate may result in slow convergence or getting stuck in a local minimum. In embodiments, learning rate is set in the range of 0.0001 and 0.1 , or between about 0.001 and 0.01 .
[0105] In Example 5, DenseNet-121 model was trained on datasets i + ii with a list of different hyperparameters (Figure 5b, Table 2). The highest validation accuracy and the lowest validation loss was obtained at using images resized at 768x768 pixels. Various batch sizes, optimizers and data augmentation techniques were explored (Figure 5b). A batch size of 16 images and the stochastic gradient descent (SGD) optimizer delivered optimal performance (Figure 5b, Table 2). The model with the best-performing parameters based on the validation set during the training stage was saved and named AINU.
[0106] Despite the relatively small number of input images (Figure 5a, datasets i + ii), AINU demonstrated exceptional proficiency in identifying somatic cells and hiPSCs, reaching an average validation accuracy of 95.98 (Figure 5b and Table 2).
[0107] AINU efficiently identifies cell types, cancer cells, and the degree of infection in cells
[0108] AINU model is shown to identify cancer cells from non-cancerous cells. AINU was trained by replacing hiPSC cells of dataset ii with HeLa cells, a type of cancer cell (Figure 2k, Figure 5a, dataset x). AINU was further retrained and validated AINU using dual-color images of H3 and Pol II and tested on the withheld dataset of dual-color nuclei images of HeLa cells (Figure 5a, dataset xi). The model performances reached a weighted accuracy and F1 score of 0.84 ± 0.09, showing effective discrimination of HeLa cells (Figure 2I, Table 11). In some embodiments, various metrics, including F1 score, validation accuracy and loss, Area Underthe Receiver Operating Characteristic curve (AUC), precision and recall were used to evaluate the model. In embodiments, the evaluation was based on the average accuracy, loss, precision or recall on k-fold cross validation.
[0109] AINU trained with SR images of DNA can correctly identify the degree of infection in HSV1-infected cells (Example 12). To evaluate the model's capacity to identify early-infected cells, lung carcinoma epithelial (A549) cells were mock-infected or infected with HSV-1 and photographed at three time points: mock, 1 hour post-infection (hpi), and 3 hpi. The result suggests that AINU correctly identified HSV-1 -infected cells already at 1 hpi, indicating the presence of subtle chromatin remodeling at 1 hpi.
[0110] Further, AINU efficiently identifies I detects preleukemic cells. AINU was retrained using RNA Pol II phSer5 super resolution images of preleukemic cells and wild type control cells (dataset xvi). The retrained model accurately classified 91% of the cases achieving a ROC AUC of 0.91 (Figure 9A, 9B). Further training with additional images including preleukemic cells with an additional mutation (dataset xvii), achieves ROC AUC greater than 0.90 (Figure 9C, Figure 9D). AINU trained with confocal images of DNA successfully detects 98% of preleukemic cells (Figures 9E, 9F) (Example 16).
[0111] The present disclosure further relates to a potentially robust diagnostic tool with multiple applications. For example, AINU may be used to identify I detect preleukemic cells in a robust manner (Example 16). Further, AINU may be used to identify hiPSC clones with high pluripotency grade, using only immunostaining instead of the current requirement for time-consuming and resource-heavy, animal-based experiments. Additionally, the methods of this disclosure may be used to identify virus-infected cells in blood or tissue at very early stages after infection, with important applications for immunology and viral biology. Finally, the methods of this disclosure may be useful in identifying cancer cells and metastatic cells among wild-type cells from human specimens.
[0112] In embodiments, images which are not super-resolution images can alternatively be used; for example, diffraction limited microscopy (e.g. confocal, widefield microscopy) can be used, e.g. in combination with cell-staining (e.g. with DAPI, Hoechst etc.), to visualise clutch-associated DNA and changes thereof.
[0113] Alternative training approaches to uncover potentially hidden nuclear features that could improve model performance
[0114] Training using alternative datasets, including a synthetic dataset that is sufficiently large (Example 1 , generating a synthetic dataset), helps training a machine learning model on diverse, representative and accurate data, and helps the model to better recognise patterns and make predictions on new, unseen data. Access to high quality data allows the users to quickly iterate and improve the models, reducing the time and resources required for development.
[0115] While the present disclosure already demonstrated high validation accuracy of 95.98 in differentiating between the various cell states using a minimal amount of training data, using dualcolor dataset consisting of 349 full-size images (Example 4), the current disclosure further provides additional datasets to uncover potentially hidden nuclear features.
[0116] In Example 1 , thousands of synthetic images of hiPSCs and somatic cells, indistinguishable from the real images by the naked eye, were generated (Figure 6c). STORM-color images representing 80% of the full size ground truth dataset were used to train and validate a StyleGAN-2 ADA Generative Adversarial Network (GAN) model using the Differentiable Augmentation method (DiffAugment).
[0117] In Example 7, patch dataset and simulated dataset were provided to improve model performance. In some examples and embodiments, patch dataset comprises of patches extracted from the dualcolor STORM images rendered at higher resolutions such as 30x instead of 10x (Figure 5a, datasets iii and iv), which are then subdivided into smaller, non-overlapping patches, increasing the size of the dataset.
[0118] Use of interpretable Al to identifying a feature associated with cellular heterogeneity
[0119] The subtlety of cellular features presents a challenge to identifying a cellular feature that accounts for cellular phenotypic heterogeneity. For example, some cellular features cannot be histologically labelled to demonstrate differences between different populations of cells. Other cellular features, such as subtle differences in morphology, requires thorough quantification.
[0120] While super resolution microscopy provides higher resolution images of cellular features, identifying the cellular features that are indicative of differences in cell identify, or labelling and classifying the cellular features, and segmenting the features are extremely challenging in practice as the distribution and differences in cellular features to be characterised are very nuanced.
[0121] Another challenge lies in the field of understanding the extracted spatial distributions and densities of molecules and structures in the cell and using this information to identify cellular features indicative of cellular heterogeneity. Further challenges lie, for example, in identifying the shapes of the clustered molecules. Although there have been various efforts to characterise spatial distributions on cells in a tissue by constructing distance matrix and calculating the nearest neighbour distance between the cells (Parra et al., 2021), such analytic methods are not suitable for identifying subcellular I cellular features indicative of cellular heterogeneity.
[0122] Therefore, it would be beneficial to have a greater understanding of the heterogeneity of cellular features using systems that extract and learn complex patterns and rules from data. The present disclosure, in aspects and embodiments, provides a method for applying Artificial Intelligence to the images to identify cellular features that are associated with cellular heterogeneity. In some aspects, identifying cellular features involved identifying image regions that were the most influential in the model’s classification decision. In some embodiments, image pixels were defined as the inputs, and the contribution of each input to the model’s output were evaluated by calculating positive or negative attribution scores for each part of the image. Positive attribution scores indicate areas of the image that contribute to the model’s prediction, while negative attribution scores highlight regions whose removal increases the probability of the predicted class. In embodiments, Interpretable Al or occlusion-based method was used to provide an understanding of the model's decisions. In some embodiments, occlusion technique, was performed by systematically sliding a 64x64 pixel and / or a 32x32 pixel window across each image, and superimposing the heatmap of attribution scores onto the original image.
[0123] Aspect and embodiments of this disclosure provide the opportunity for Interpretable Al to identify Pol II localizations within the nucleoli as the key feature characterising hiPSCs. In an example, occlusion-based method identified the nucleoli as the predominant positive attributions for hiPSCs, with emphasis on specific areas within the nucleoli. (Figures 8b-8d). In some embodiments, class activation mapping (CAM) technique can be applied to validate the area of the image or the feature that contributed to identify cellular heterogeneity. In an example, use of CAM consistently highlighted the nucleoli as the most important feature for hiPSC classification (Figure 4b, Figure 4c). Based on this result, the researchers were further able to validate that somatic cells have higher median clusters areas of Pol II, and higher median localizations per cluster in nucleoli than hiPSCs (Figure 8h), and further identify that the significantly higher expression level of asincRNAs in hiPSCs (Figure 4k). In another example, the Grad-CAM method identifies regions of preleukemic cells that contribute to relevant features leading to classification (Figure 12 and Example 18)
[0124] Interpretable Al was used to discover the decisions made by the model, which provided information about the differences in identified nanostructures. The AINU method was able to identify hiPSCs by pinpointing RNA Pol II localizations within the nucleoli, a feature that emerges from microscopy - e.g. diffraction limited microscopy, or more particularly from super-resolution microscopy data. Overall, AINU can serve as a powerful tool for studying cell heterogeneity, starting from SMLM imaging of the nucleus, and can become a potent diagnostic method.
[0125] Example 1. Image Acquisition and sample preparation for somatic cells and hiPSCs.
[0126] Datasets
[0127] 349 dual-color super resolution labelled images of core histone 3 (H3) and RNA Polymerase II (Pol II) are imaged in nuclei from human somatic cells (171 images) or hiPSCs (178 images). 384 single-color images of H3 or Pol II cells are obtained from somatic cells (187 images) or hiPSCs (197 images) (Figure 5a, Table 1). Somatic cells comprised different cell types (B lymphocytes, bone marrow mesenchymal stem cells, Muller glia, myocardiocytes, BJ fibroblasts, ARPE-19, urine epithelial cells), and hiPSCs were generated from some of the somatic cells (Table 1).
[0128] Table 1. Sources of somatic and hiPSCs obtained. Each category represents different cell types.
[0129] Separately, in nuclei of HeLa cells, single-color images of Pol II (33 images), dual-color images of H3 and Pol II (30 images), single-color images of DNA in human somatic cells (68 images), singlecolor images of DNA in hiPSCs (117 images) are obtained. Images of DNA in nuclei of mock- (81 images) and HSV1 -infected A549 cells at 1 hpi (98 images) and 3 hpi (85 images) are obtained. To avoid sex chromosome bias during model training images were collected from both male and female cells.
[0130] Cells and culture conditions
[0131] The following cells were used in this study (all human except the monkey cells): A549 cells (lung carcinoma, American Type Culture Collection, ATCC CRL-185); B lymphocytes (GM12878); BJ fibroblasts (human foreskin skin fibroblasts) (American Type Culture Collection, ATCC CRL-2522); IMR90 (fibroblasts isolate from lung tissue, American Type Culture Collection, ATCC CCL-186); bone marrow mesenchymal stem cells (MSC); HeLa; Muller glia; myocardiocytes; spontaneously arising retinal pigment epithelium (ARPE-19); urine epithelial cells; and Vero (African green monkey) cells. Further, hiPSCs were induced from amniocytes, BJ fibroblasts, dermal fibroblasts, normal lung tissue fibroblasts (hiPS(IMR90)-4: WiCell, #WISCi004) periosteum cells, umbilical cord mesenchymal stem cells (UCMSCs), and urine epithelial cells. HeLa, BJ fibroblasts, IMR90 fibroblasts and myocardiocytes were cultured in Dulbecco’s Modified Eagle Medium (DMEM) (Gibco, #C11995500BT) supplemented with 10% fetal bovine serum (FBS) (Gibco, #10099141) and 1% penicillin-streptomycin (pen / strep) (Gibco, #15140122). Urine epithelial cells were cultured in Lonza REGM BulletKit (Lonza, #CC-3190) supplemented with 1% pen / strep. B lymphocytes were cultured in RPMI 1640 with L-glutamine (Corning Cellgro, #10-040-CVR) supplemented with 15% FBS and 1% pen / strep. Muller glia and ARPE-19 cells were cultured in DMEM / F-12, GlutaMAX™ (Gibco, #10565018) supplemented with 10% FBS and 1% pen / strep. Bone marrow mesenchymal stem cells were cultured in DMEM supplemented with 0.01% human FGF-basic (Peprotech, #100-18B), 10% FBS and 1 % pen / strep. All hiPSCs were cultured in mTeSR™1 (Stem Cell, #85850) supplemented with 1% pen / strep. A549 cells were culture in F-12K (Kaighn’s) Medium (Gibco, Thermo Fisher Scientific, #21127022) supplemented with 10% FBS (Thermo Fisher Scientific, #10270106), 1 % pen / strep. Vero cells were culture in DMEM, high glucose, GlutaMAX™ Supplement (Thermo Fisher Scientific, #10566016) supplemented with 5% FBS, 1 % pen / strep.
[0132] Immunolabeling for super-resolution imaging (STORM)
[0133] Cells were plated in chambered coverglass (p-Slide 8 Well Chamber Slide, Ibidi, #80826) at a concentration of 40,000 cells / cm2 in growth medium for 24 h. Cells were fixed with 4% PFA (Leagene, #DF10315) for 15 min at room temperature and then rinsed with PBS three times for 5 min each. Cells were permeabilized with 0.4% (v / v) Triton X-100 (Across Organics, 327371000) in PBS for 15 min and then blocked in blocking buffer [10% bovine serum albumin (BSA), #B2064- 50G; 0.01 % (v / v) Triton X-100 in PBS] for 1 h at room temperature. Cells were incubated with primary antibodies [anti-RNA polymerase II CTD repeat YSPTSPS (hosphor S5) antibody, Abeam, #ab5131 ; Histone H3 antibody, Active Motif, #39763] in blocking buffer at 1 :50 dilution for STORM of both single-color imaging and dual-color imaging, overnight at 4 °C. On the second day, cells were washed three times for 5 min each with wash buffer (2% BSA and 0.01% Triton X-100 in PBS) and incubated in secondary antibody. For single-color imaging of the Pol II, the secondary antibody [donkey anti-rabbit IgG H&L (Alexa Fluor® 647), Abeam, #ab150075] was added at 1 :200 dilution in blocking buffer for 1 h at room temperature. For dual-color imaging, homemade38 dye pair [Alexa Fluor™ 405 NHS Ester, Thermo Fisher, #A30000; Cy®3 Maleimide Mono-Reactive Dye Pack, Sigma-Aldrich, #GEPA23031 ; Alexa Fluor™ 647 NHS Ester (succinimidyl ester), ThermoScientific, #A20006] labeled secondary antibodies [AffiniPure donkey anti-rabbit IgG (H+L), Jackson ImmunoResearch, #711005152; Peroxidase AffiniPure goat anti-mouse IgG, Fey subclass 1 specific, Jackson ImmunoResearch, #115-005-205] were added at a 1 :50 dilution in blocking buffer and were incubated for 1 h at room temperature. Cells were then washed three times for 5 min each with wash buffer, and once with PBS. Samples were stored in PBS at 4 °C prior to imaging.
[0134] For Figure 8h, anti-RNA polymerase II CTD repeat YSPTSPS (phosphor S5; Abeam #ab5408) was used for STORM, and anti-NOLC1 (Abeam #ab184550) was used to mark nucleoli (both diluted 1 :50). Secondary antibodies (goat anti-mouse IgG Alexa Fluor™ 647, ThermoScientific #A21235; and goat anti-rabbit IgG, Oregon Green 488, ThermoScientific #0-11038) were used at a 1 :250 dilution. Immunostaining was performed using the protocol given above. The complete STORM protocol can be found in Martin et al.
[0135] STORM imaging
[0136] STORM imaging was performed: i) for proteins, using a N-STORM microscope (Nikon) equipped with a CFI SR HP Apochromat TIRF 100XAC oil objective and an ORCA-Flash 4.0 V3 Digital CMOS Camera; ii) for DNA, using ONI (Oxford Nanoimager); and iii) for Pol II and conventional for N0LC1 (Figure 8h), using a N-STORM 4.0 microscope (Nikon) with an iXon Ultra 897 camera (Andor) and a CFI HP Apochromat TIRF *100 1.49 oil objective, under highly inclined and laminated optical sheet (HILO) illumination mode. NIS element (4.60 and 5.21) software was used for NSTORM image acquisition. Only nuclei that were morphologically coherent with the characteristics expected of an interphase nucleus were imaged.
[0137] For single-color STORM imaging of Pol II, continuous imaging acquisition was performed with simultaneous 405 and 647 nm illumination of the sample, 10 ms exposure time for 60000 frames. 647 nm laser was used at constant 3.5 kW / cm2, and 405 nm laser power was gradually increased over the imaging.
[0138] For dual-color STORM imaging of Pol II and H3, a double activator and single reporter strategy was used by combining AF405_AF647 anti-mouse and Cy3_AF647 anti-rabbit secondary antibodies. Sequential imaging acquisition was performed (1 frame of 405 nm activation, followed by 4 frames of 647 nm reporter; and 1 frame of 560 nm activation, followed by 4 frames of 647 nm reporter) with 20 ms exposure time for 120,000 frames. The 647-nm laser was used at a constant 3.5 kW / cm2 power density, while the 405-nm and 560-nm laser powers were gradually increased over the imaging.
[0139] Single-color imaging of DNA images was performed on the NanoimagerS Mark II from ONI (Oxford Nanoimaging) equipped with an Olympus 1.4NA 100* oil immersion super apochromatic objective and a Hamamatsu sCMOS Orcaflash 4 V3. Continuous imaging acquisition was performed with simultaneous 647 nm illumination of the sample, 10 ms exposure time for 45,000 frames.
[0140] During the super resolution imaging, freshly prepared buffer was added to the sample: 100 mM cysteamine MEA (Sigma-Aldrich, #30070), 1 % Glox solution (0.5 mg / ml glucose oxidase, 40 mg / ml catalase; Sigma-Aldrich, #G2133 and #C100), 5% glucose (Sigma-Aldrich, #G8270) in PBS. The buffer was exchanged every hour to avoid enzimatic activity decay and maximize the signal-to- noise ratio of the localizations.
[0141] For ONI images, localization lists were obtained using the integrated simultaneous localization software from ONI (ONI NimOS v.10.5) and subsequently converted into lnsight340 compatible files for post-processing using a custom-built software in MATLAB 2016a.
[0142] N-STORM and ONI microscopes have a resolution of 160nm / pixel and 117nm / pixel respectively. N-STORM images were analyzed and rendered in lnsight340 as described38,41. Specifically, point spread functions (PSFs) from the emission from single fluorophores were identified at individual frames of the acquired videos, based on a restrictive set thresholds of intensity, PSF size and symmetry. They were fit to a two-dimensional Gaussian; from which the centroid of the PSF was obtained and the x and y positions for 2D STORM imaging were then obtained from the PSF. All images were acquired using HiLo to maximize the signal-to-noise ratio in the nucleus. Focus plane was set to the z of maximus nucleus area to ensure the representativity of the slices. Cells with uneven intensity profiles or localization artifacts were manually discarded before proceeding with the DL pipeline.
[0143] For images acquired using the ONI Nanoimager with two different spectrally separate fluorophores, a prior calibration using tetraspeck beads was performed according to the instructions of the manufacturer, to ensure a minimal drift (around 5-20 nm) between channels.
[0144] EdC incorporation and DNA labeling
[0145] To label DNA, a 48-h incorporation of 5 pM EdC (5-ethynyl-2'-deoxycytidine; Sigma-Aldrich, #T511307) was performed using the following cells: HeLa, BJ fibroblast, myocardiocytes, urine epithelial cells, B lymphocytes, Muller glia, ARPE-19, MSCs, and hiPSCs induced from amniocytes, BJ fibroblasts, dermal fibroblasts, periosteum cells, UCMSCs and urine epithelial cells. Cells were plated in chambered coverglass (p-Slide 8 Well Chamber Slide, Ibidi, #80826) at a concentration of 40,000 cells / cm2 in growth medium supplemented with 5pM EdC for 24 h. For somatic cells, the growth medium was replaced on the second day for resting medium supplemented with EdC and incubated another 24 h. For hiPSCs, 10 pM Rock inhibitor (Rock inhibitor Y-27632, Millipore, SCM075) was added together with 5 mM EdC for the first 24 h, then was replaced for medium supplemented with 5 mM EdC only. Cells were fixed with 4% PFA (Leagene, #DF10315) for 15 min at room temperature and then rinsed with PBS three times for 5 min each. Cells were permeabilized with 0.4% T riton X-100 in PBS for 15 min and rinsed with PBS three times for 5 min each. The click chemistry reaction was performed by incubating cells for 45 min at room temperature in click chemistry buffer [150 mM HEPES pH 8.2 (Sigma-Aldrich, #7365-45-9), 50 mM amino guanidine (Sigma-Aldrich, #396494), 100 mM L-ascorbic acid (Sigma-Aldrich, #A92902), 1 mM CuSO4 (Sigma-Aldrich, #C1297), 2% glucose (Sigma-Aldrich, #G8270), 0.1% Glox solution (described above, in STORM imaging), and 10 mM AF647 azide (Thermo Fisher Scientific, #A10277)], protected from light. After three washes with PBS, STORM imaging was directly carried out for single-color DNA imaging experiments (see STORM imaging).
[0146] Generating simulated images
[0147] To explore alternative training approaches, a simulated dataset was generated with DiffAugment (Figure 5a, dataset v), to expand both the training and validation sets. During the GAN training process, authentic SR dual-color images from the train and validation datasets (Figure 5a, dataset i) were used to train a model capable of generating images indistinguishable from the real images by the naked eye (Figure 5a, dataset v). The SR dual-color images consist of a sample of 290 10x-rendered images from dual-color STORM localizations (representing 80% of the whole dataset, out of a total of 349). These were uniformly resized to 1024x1024px, and used to train and validate a StyleGAN-2 ADA Generative Adversarial Network (GAN) model using the Differentiable Augmentation method (DiffAugment). Default ‘color, translation, cutout’ augmentation technique was applied with kimg option set to 500. The purpose of this model was to generate thousands of synthetic images of hiPSCs and somatic cells, with 5000 images of each condition generated (Figure 6c). The remaining 59 images (representing 20% of the whole dataset, out of the total of 349) from dual-color STORM localizations were set aside and used later as a test set for the best performing CNN model trained with synthetic images. The quality of the simulated images was assessed using the common Frechet Inception Distance (FID) reaching high values of 40.7 for somatic and 37.5 for hiPSC cells.
[0148] Example 2. Pre-processing somatic cell and iPSC images
[0149] The molecule localization coordinates of the selected region of interest (ROI, i.e., nucleus) of dualcolor Pol II and H3 were multiplied either by 10 or 30 corresponding to a resolution of 16 nm / px and 5.3 nm / px respectively and then rendered into images by computing a Gaussian kernel density estimate using Scikit-learn KDTree function. The images can be resized at the desired dimension to be used as an input data while keeping the aspect ratio.
[0150] To enable the model to differentiate between the background of the image outside the nucleus and regions within the nucleus lacking localizations, the background of full-size 10x rendered images are set to white.
[0151] Both single and dual-color images rendered at 10x magnification (Figure 5a) were directly used as train / validation and test sets. Dual-color images rendered at 30x magnification were divided in nonoverlapping patches (Figure 6a) of 3 different sizes (256x256, 512x512, 768x768 pixels). Only patches with more than 75% of the pixels containing localizations were used as training / validation and test sets.
[0152] To avoid changing the aspect ratio, which would have modified the spatial relationship of the imaged molecules, each image was first resized on its longest axis to the desired value (e.g., 1024 pixels for the simulation process) by keeping the aspect ratio and was subsequently padded with background color on the shortest axis. For all other datasets (i.e., single-color Pol II, DNA and dual-color Pol II and H3 HeLa) 10x magnification images of the molecule localizations were rendered and used.
[0153] Example 3. Training a CNN model to identify somatic cells and hiPSCs
[0154] A CNN model was trained on chromatin clutch structure or on RNA Polymerase II (Pol II) nuclear distribution. Using dual-color STORM, the inventors co-imaged the nucleosome core histone H3 and Pol II in 349 nuclei of human somatic cells and hiPSCs obtained from different somatic cell types (Table 1). The super resolution localizations of the two targeted molecules were used to generate the full-size ground truth dataset, with images rendered at 10x magnification (Figure 1 a and Figure 5a), partitioned into training, validation and withheld test sets.
[0155] Hardware and software All computational operations were conducted using the Center for Genomic Regulation cluster infrastructure, which comprises 10 NVIDIA GeForce 2080 Ti with 11 Gb of memory per GPU. The DL computations were performed using the PyTorch framework version 1.10, with CUDA drivers version 11.3 and Python version 3.9.7. For ONI images, molecule localizations were extracted and exported with ONI NimOS v.10.5 and then converted into insights compatible files for post processing with MATLAB 2016a. For N-STORM images, molecule localizations were extracted using Insights. The images were segmented by manually selecting the nuclear areas corresponding to each cell nucleus with Fiji. Plots and statistical tests of Fig. 4i-k were produced with R version 4.2.2 and GraphPad Prism version 8.01.
[0156] Example 4. Identifying the optimal machine learning architecture
[0157] To identify the optimal machine learning architecture and its hyperparameters, a stratified 5-fold cross-validation (CV) approach was used, partitioning the full-size ground truth dataset with a 80 / 20 training and validation set ratio. 11 distinct CNN architectures were selected based on unique attributes and suitability for the cell classification task, including several models that each offered a different approach and computational efficiency (Figure 5a, Table 2).
[0158] The input dataset used was dual-color images of core histone 3 (H3) and RNA Polymerase II (Pol II). This dual-color dataset, consisting of 349 full-size images rendered at 10x magnification, underwent an 80 / 20 split into training and validation sets, keeping the proportion of each cell class present in the dataset. Employing a stratified 5-fold cross-validation, 11 different CNN architectures were each trained for 300 epochs using images resized to 512x512 pixels. For all 11 models trained, hyperparameters remained constant for fair comparisons (Table 3).
[0159] The best-performing architecture was chosen based on the average accuracy and loss over the five folds. DenseNet-121 was the top performing CNN architecture in identifying somatic cells and hiPSCs, with an average validation accuracy of 92.26 and an average loss of 0.292 (Figure 5b, Table 2), and was used for subsequent analyses. able 2. List of CNN architecture, image size, batch size, optimizer and data augmentation technique and learning rate tested. Top performers in each category are in bold.
[0160] Table 3. List of Hyperparameters used for training CNN architectures.
[0161] Example 5. Identifying image size, batch size, optimizer and data augmentation technique for the machine learning model architecture
[0162] To find the most suitable common image size as a network input, DenseNet-121 model was trained and validated using images (Figure 5a, datasets i + ii) rescaled to different pixel sizes (Table 2). The highest validation accuracy and the lowest validation loss was obtained by using images resized at 768x768 pixels.
[0163] Data augmentation techniques, including flipping, random rotation, and color adjustments, were applied using the Albumentations library (Figure 5a, Table 2). Region dropout and the Grid Mask technique were explored to enhance training data diversity (Figure 5b, Table 2). The use of the Grid Mask technique with a ratio of mask holes to unit size set at 0.3 and grid unit sizes ranging from 128 to 384 led to the highest average accuracy and lowest average loss values when assessed through 5-fold cross-validation (Figure 5b, Table 2). Various batch sizes, optimizers and data augmentation techniques were explored (Figure 5b). A batch size of 16 images and the stochastic gradient descent (SGD) optimizer delivered optimal performance (Figure 5b and Table 2).
[0164] The model with the best-performing parameters (Figure 5b, Table 2) based on the validation set during the training stage was saved and named AINU.
[0165] Despite the relatively small number of input images (Figure 5a, datasets i + ii), AINU demonstrated exceptional proficiency in identifying somatic cells and hiPSCs, reaching an average validation accuracy of 95.98 (Figure 5b and Table 2).
[0166] Example 6. Identifying human somatic cells and hiPSCs using dual-color images of chromatin features
[0167] AINU's final architecture and optimal hyperparameters were used for training, validation, and evaluation across various datasets, including single-color images of Pol II localizations collected in somatic cells and hiPSCs, dual-color images of Pol II and H3 collected from HeLa cells, and single-color images of DNA collected from somatic cells, hiPSCs and HSV1 -infected A549 cells. The model that exhibited the best performance during the 5-fold CV underwent a retraining and validation process lasting 300 epochs. This retraining phase involved randomly dividing the whole dataset into training, validation and withheld sets, with 20% in the withheld set (Figure 5a, datasets i and ii). Training, validation and evaluation on a withheld test set was repeated five times to mitigate any potential biases stemming from random dataset splitting or image heterogeneity. On average, the model achieved a weighted accuracy and F1 score of 0.85 with a standard deviation (SD) of 0.07, and a ROC AUC score of 0.95 ± 0.04 (Table 4). In the best dataset split, the model excelled further, achieving a weighted accuracy and F1 score of 0.94 (Figure 1 b, Table 4), a ROC AUC score of 0.98 (Figure 1c), and impressive average precision and recall scores of 0.98 and 0.99 for somatic cells and hiPSCs, respectively (Figure 1 d, Figure 1e).
[0168] Table 4. Precision, Recall, F1-score and Support values from each of 5-fo d cross validation.
[0169] Example 7. Uncovering potentially hidden nuclear features by training AINU on patch dataset and simulated dataset
[0170] To increase the training dataset and to uncover potentially hidden nuclear features that could improve model performance, AINU was trained on a patch dataset, using patches extracted from the dual-color STORM images rendered at higher resolutions rendered at 30x magnification (Figure 5a dataset iii, Figure 6a). For consistency with previous experiments, AINU was trained for 300 epochs and selected the best models based on their corresponding validation set performance. To collect patches (Figure 5a, datasets iii and iv), the localization coordinates of each blinking fluorophore were multiplied by a factor of 30, effectively rendering images at a 30x magnification relative to the original camera pixel size. These high-resolution images were then subdivided into smaller, non-overlapping patches, of 256x256 pixels or 512x512 pixels, effectively increasing the size of the data set.
[0171] The performances of the AINU models independently trained with patched or simulated images were evaluated using separate withheld test sets that are selected randomly (Figure 5a, dataset iv). This test set is different from those used with full-size images (Figure 5a, dataset ii). The evaluation included using various metrics, including F1 score, validation accuracy and loss, Area Under the Receiver Operating Characteristic curve (AUC), precision and recall. A simulated dataset involved generating a large set of synthetic images indistinguishable from the real images generated using a Generative Adversarial Network (GAN) model (Example 1 , Generating a simulated dataset). When utilizing simulated images, AINU was trained and validated using 4000 and 1000 simulated images, respectively, and was evaluated on a set of 59 full-size 10x-rendered real images that were previously kept separate during the training of the GAN model for synthetic image generation. AINU reached a modest weighted accuracy and F1 score of 0.78 ± 0.11 at a 95% confidence interval (Cl) (Figures 6d-6g, Table 5). For the GAN- simulated images, AINU yielded a very high validation performance using images of 512x512 pixel size; however, it performed less well (but still better than random) in its withheld test set of real images (Figure 5a, dataset ii). It is possible that the reduced ability of the GAN technique to partially capture the diversity of unseen images may be due to the use of a validation set of simulated images, which is too close to the distribution of training images and, thus, may lead the model to overfit the training data.
[0172] Table 5. Precision, Recall, F1-score and Support values from AINU trained with simulated dataset.
[0173] Due to the intrinsic differences between full-size dual-color images, patch-based images, and simulated images, fine-tuning of AINU was needed when training with patches or simulated images. When evaluating the model trained with patches, a classification threshold of 50% was applied as the minimum number of patches correctly assigned to the positive class. This means that an image was correctly predicted if more than 50% of its patches were assigned to the true class. No transformations that altered the aspect ratio of the image were applied to augment the data, as this would have resulted in a distorted spatial relationship of the imaged molecules, which is a key characteristic feature of somatic and hiPSCs. Hence, image resizing to 1024x1024px was performed.
[0174] Because both models trained with simulated images or patched images were not trained five times with random dataset splits due to high resource consumption, the performances scores were reported with the 95% Cl calculated with the normal approximation method.
[0175] When training was performed using patches, AINU exhibited optimal validation performance when using non-overlapping patches of 256x256. However, when challenged with a test set of real images (Figure 5b, dataset iv) it showed a lower performance than the original training strategy (full-size image), with a weighted accuracy and F1 score of 0.8 ± 0.09 (Figures 6h-k, Table 6). This suggests that the use of only partial sections of the nucleus, even if at higher magnification, does not lead to improved overall performance.
[0176] Table 6. Precision, Recall, F -score and Support values from AINU trained with patch dataset.
[0177] Finally, to validate and compare the results of AINU obtained using the three different approaches (full-size images - dataset i, simulated images - dataset v and patches - dataset iii) under the same conditions, the researchers retrained AINU with the three datasets but using the same exact splits and withheld (unseen images) test set. Specifically, the same test set as for the GAN simulated dataset (Figure 5a, dataset ii) was used. The performance of the three models matched those obtained previously on different test sets, whereby the model trained with full-size images had the best performance, with a weighted accuracy of 0.85 ± 0.09 (Table 7).
[0178] Table 7. Precision, Recall, F1-score and Support values from AINU trained with full-size images (dataset i), simulated images (dataset v) and patches (dataset iii).
[0179] Example 8. Performance of AINU trained on specific cell types
[0180] To test the performance of AINU on unseen images from unseen cell types, AINU was retrained using full-size real images and 5-fold cross validation (Figure 5a, dataset i) but withholding all images of urine epithelial cells or BJ fibroblasts. AINU still performed well: with a 95% Cl, it gave a weighted accuracy and F1 score of 0.8 ± 0.11 when urine epithelial cells were withheld, and 0.79 ± 0.09 when BJ fibroblasts were withheld (Figures 7a-h, Table 8). Then the model was trained by withholding different cell types that are present in only one state (either hiPSC or somatic). For instance, when amniocytes-derived hiPSCs and human mesenchymal stem cell (hMSC) types were withheld, the model still achieved above-average performances, with a weighted accuracy and F1 score of 0.69 ± 0.12 (Table 8). These results suggested that the spatial organization of the imaged molecules was better determined by the cell state (somatic vs hiPSCs) than the cell type, and cell states are better captured when the cell types that are withheld from training and used as a test set belong to the same type.
[0181] Table 8. Precision, Recall, F1 -score and Support values from AINU trained with full-size dataset (Figure 5a, dataset i) excluding BJ fibroblasts.
[0182] Example 9. AINU trained with Pol II localizations correctly identifies the cell type As dual-color STORM imaging is time-consuming and technically challenging, single-color superresolution imaging was used for training. Datasets were generated by rendering images from the Pol II channel within the dual-color SR localizations (Figure 2a and Figure 5a, datasets vi and vii) and added 384 images from single-color Pol II localizations (Figure 2f and Figure 5a, datasets viii and ix). Five cycles of retraining, validation, and testing of AINU, each time using random dataset splitting of 733 Pol II images were performed. The model's performance closely paralleled that of the model trained with dual-color images, with an average weighted accuracy of 0.87 ± 0.03 and an average AUC score of 0.95 ± 0.01 (Table 9). In the best dataset split AINU reached an accuracy and F1 score of 0.93 for hiPSCs and 0.88 for somatic cells (Figure 2b), an AUC score on the withheld set of 0.97 (Figure 2c), and an average precision and recall of 0.97 for both somatic cells and hiPSCs (Figure 2d, Figure 2e).
[0183] Table 9. Precision, Recall, F1-score and Support values from AINU trained with single-color Pol II dataset (Figure 5a, datasets vi and vii).
[0184] Then AINU model was trained for 300 epochs using only the 384 images of single-color Pol II localizations for five times, with random splitting each time. The model performed slightly better with this smaller dataset, with an average weighted accuracy and F1 score of 0.91 ± 0.05, and an average AUC score of 0.97 ± 0.02 (Table 10). Further, it achieved remarkable performance metrics during the best round of training, validation and splitting (Table 10), with 0.94 for hiPSCs and 0.97 for somatic cells (Figure 2g), an AUC score of 0.99 (Figure 2h) and an average precision and recall of 0.99 for both somatic cells (Figure 2i) and hiPSCs (Figure 2j). The increased performance could be due to a higher distribution consistency of the single-color dataset images (which were from the same experiment) than that of the dual-color Pol II images. The distribution of Pol II localizations was the primary feature learned by the model to achieve its performance irrespective of whether dual-color or single-color images were used.
[0185] Table 10. Precision, Recall, F1-score and Support values from AINU trained with smaller singlecolor Pol II dataset (Figure 5a, datasets vi and vii).
[0186] Example 10. AINU trained with Pol II signatures could identify cancer cells
[0187] AINU model was trained using the same withheld test set of 39 previously unseen somatic images (Figure 5a, dataset ii, somatic cells only) but replacing the hiPSC images with 33 single-color, full- size rendered images of Pol II localizations collected in HeLa cells, a type of cancer cell (Figure 2k, Figure 5a, dataset x). A weighted accuracy is approximately 0.78 ± 0.1 95% Cl, with 0.58 ± 0.17 for HeLa and 0.95 ± 0.07 for somatic (Table 11).
[0188] Given that the accuracy in identifying HeLa cells only slightly exceeded (>8%) random chance, AINU was retrained and validated by including six HeLa nuclei images in the training dataset, and two in the validation dataset. The performance of the model with a withheld test set of unseen HeLa images increased significantly, reaching a weighted accuracy of 0.84 ± 0.09, with 0.72 ± 0.18 for HeLa and 0.92 ± 0.08 for somatic cells (Table 11). Then the same strategy was applied to retrain and validate AINU using dual-color images of H3 and Pol II and tested on the withheld dataset of dual-color nuclei images of HeLa cells (Figure 5a, dataset xi). The model performances increased further, reaching a weighted accuracy and F1 score of 0.84 ± 0.09 (Figure 2I, Table 11), an AUC score of 0.95 (Figure 2m) and an average precision and recall of 0.97 for somatic cells (Figure 2n) and 0.92 for HeLa cells (Figure 2o). Thus, the AINU model trained with hiPSCs and somatic cells could be used to detect cancer cells after appropriate retraining with a few additional cancer cell images, opening up future possible applications for AINU.
[0189] Table 11. Precision, Recall, F1 -score and Support values from AINU retrained and validated by six HeLa nuclei images in the training dataset, and two in the validation dataset (Figure 5a, dataset x).
[0190] Example 11. AINU trained with SR images of DNA can correctly identify hiPSCs, somatic cells and HSV1 -infected cells
[0191] To test whether the same model trained with SR images of DNA could also yield accurate results, SR DNA imaging via click chemistry was performed. A dataset of 185 single-color SR images of DNA (Figure 5a, datasets xii and xiii) in human somatic cells (Figure 3a) and hiPSCs (Figure 3b) was generated, using 10-times higher resolution and resizing to 768x768 pixels but maintaining the aspect ratio through padding techniques. The model was trained for 300 epochs, with five random splits of the dataset and consistent use of the hyperparameters and augmentation techniques previously used for the full-size dual-color images. On a withheld test set (Figure 5a, dataset xiii), the model resulted in an average weighted accuracy of 0.94 ± 0.04 and average AUC score of 0.98 ± 0.03 (Table 12). On the best dataset random split, the model achieved perfect performances, with a weighted F1 score, AUC and accuracy of 1 for both hiPSCs and somatic cells (Table 12, Figure 3c, Figure 3d).
[0192] Table 12. Precision, Recall, F1-score and Support values from AINU trained using SR DNA imaging via click chemistry (Figure 5a, dataset xii and xiii).
[0193] Example 12. AINU trained with SR images of DNA can correctly identify the degree of infection in HSV1 -infected cells
[0194] The nuclei of HSV-1 infected cells reorganize to form replication compartments after infection. To test the model's ability to identify early-infected cells, lung carcinoma epithelial (A549) cells were mock-infected or HSV-1 infected and then imaged at three different time points: mock, 1 hour postinfection (hpi) and 3 hpi. Remodeling of the host chromatin fibers was visually evident as early as 3 hpi (Figures 3e-3g). The model was trained, validated and tested on a withheld test set five times with different random splits of the whole dataset (Figure 5a, datasets xiv and xv). Despite the low number of input images (264), the model performed well on the withheld test set (Figure 5a, dataset xv), achieving an average weighted accuracy of 0.83 ± 0.03 and micro and macro average ROC AUC scores of 0.95 ± 0.01 , over five runs with random splits (Table 13). With the best dataset split, the model reached a weighted accuracy and F1 score of 0.87 (Table 13), with individual accuracies of 0.69, 0.9 and 1 for mock, 1 hpi and 3 hpi cells, respectively (Figure 3h), and micro and macro average ROC AUC scores of 0.98 (Figure 3i). This suggests that AINU correctly identified HSV-1 -infected cells already at 1 hpi, indicating the presence of subtle chromatin remodeling at 1 hpi. Using DNA SMLM to train a CNN model is a viable option, and that the model can correctly distinguish different levels of chromatin organization.
[0195] Table 13. Precision, Recall, F1 -score and Support values from AINU trained using whole HSV dataset (Figure 5a, datasets xiv and xv).
[0196] Example 13. Interpretable Al identifies the nucleoli as cell discriminator
[0197] To evaluate the contribution of each input feature (image pixels) on the model’s output and prediction accuracy and to gain an understanding of the model's decisions, the Captum Framework and its occlusion-based method were used. Captum framework for PyTorch was used to analyze the key features behind the model’s classification, and to explain the relationship between the spatial organization of molecules and their impact on biological functions.
[0198] This approach allowed to visually inspect the areas of the image that exhibited positive or negative attribution scores for each correctly predicted sample (Figure 8a). Positive attribution scores indicate areas of the image that contribute to the model’s prediction, while negative attribution scores highlight regions whose removal increases the probability of the predicted class. To apply the occlusion technique, a 64x64 pixel window, with a 32-pixel stride, was slid across each image generated from Pol II localizations.
[0199] Occlusion-based algorithm (Figure 4a) was first used to visually inspect the areas of the image that exhibited attribution scores that were positive (contributing) or negative (detracting) for each correctly predicted sample, using a 64x64 occluding pixel window, with a 32-pixel stride. Notably, the algorithm assigned the highest positive scores (green) to regions corresponding to nucleoli in hiPSCs, and negative scores (red) to regions at the nucleus’s edge or interior (Figure 4a, upper panel). Conversely, somatic cells showed the opposite pattern, with positive scores primarily in the nuclear periphery and negative scores within or near the nucleoli (Figure 4a, bottom panel).
[0200] Then 32x32 pixel occluding windows with different strides (e.g.,16 or 8 pixels) were used for a finer resolution. Consistently with the previous 64x64 / 32 data (Figure 8b), the algorithm still identified the nucleoli as the predominant positive attributions for hiPSCs (Figures 8c, 8d), with emphasis on specific areas within the nucleoli.
[0201] Subsequently, the heatmap of attribution scores was superimposed onto the original image, facilitating the identification of the image regions most influential in the model’s classification decision.
[0202] The result was confirmed by applying the class activation mapping (CAM) technique to each correctly predicted image, which consistently highlighted the nucleoli as the most important feature for hiPSC classification (compare Figures 4b-c, using CAM, to Figure 4a, with the occlusion-based method of the same image). The differences in localization quantification, distribution and organization inside the nucleoli of hiPSCs and somatic cells were verified using a cluster analysis method previously developed from the researchers (Figures 4i,j).
[0203] Using the occlusion-based algorithm and the CAM technique gave consistent results also with H3 and Pol II dual-color images, with positive values assigned to the nucleoli for hiPSCs and to the outer region of the nucleus for somatic cells (Figure 4d). Likewise, using the occlusionbased method on images of hiPSCs and somatic cells from single-color localization of DNA identified the edge of the nucleus as primarily positive in somatic cells and predominantly negative in hiPSCs (Figure 8e).
[0204] To validate the importance of nucleoli for cell state discrimination, the test set images were modified by replacing the nucleoli regions with cloned areas from the corresponding image (Figures 8f, 8g). The accuracy of identification was significantly decreased for hiPSCs, for which nucleoli are a crucial distinguishing feature, and slightly increased for somatic cells, in which nucleoli were considered ‘distracting’ areas (e.g., with negative attribution scores) (Figures 4e-h, Table 14). These results emphasized the importance of nucleoli for classifying hiPSCs, and the ability of the model to accurately identify these key structures. able 14. Full-size dual-color images of H3 and RNA Pol II with occluded nucleoli.
[0205] Pol II in the nucleolus transcribes antisense intergenic noncoding RNAs (asincRNAs) from rDNA large intergenic spacers (IGSs) thereby controlling RNA polymerase l-derived sincRNAs. Hence the researchers analyzed whether the amount or spatial organization of the Pol II localization signals differed in the nucleoli of 13 hiPSC and 12 somatic cells stained with nucleolar and coiled-body phosphoprotein 1 (NOLC1) (Figure 8h). The inventors’ previously developed clustering algorithm (Ricci et al., 2015) revealed that somatic cells have higher median clusters areas of Pol II, and higher median localizations per cluster in nucleoli, than hiPSCs, thus supporting our findings from the occlusion-based and CAM methods (Figure 4i, Figure 4j).
[0206] Finally, using single-strand reverse transcription quantitative PCR (ssRT-qPCR) in IMR90- fibroblast somatic cells and IMR90-derived hiPSCs, it was observed that there is a significant increase in the expression levels of the asincRNAs transcribed from the rDNA large IGS28 in hiPSCs as compared to somatic cells (Figure 4k). Overall, these results show that AINU can discriminate hiPSCs from somatic cells by identifying Pol II localizations and activity in the nucleoli of the two cell states. ssRT-qPCR amplification of asincRNAs RNA extraction hiPSCs (generated from IMR90, WiCell, no. WISCi004) and IMR90 cells at 70% to 80% confluency were washed with Rnase-free PBS before RNA isolation using a Qiagen Rneasy mini Kit (Qiagen, #74106). The isolated RNA was treated with Dnase I enzyme (Thermo Fisher Scientific, #18068015) and precipitated with 0.8 M trisodium citrate and 1.2 M NaCI.
[0207] Strand-specific cDNA synthesis
[0208] IGS region 28 was chosen based on being among the most significant transcriptional loci identified in Abraham et al. (2020) Nucleolar RNA polymerase II drives ribosome biogenesis, Nature 585, 298-302. The cDNA synthesis reaction was set up for the transcript of interest using M-MLV reverse transcriptase (Thermo Fisher Scientific, #28025013). 10 pM of one FS-A primerwas used, plus 7SK-A as control (Table 15). The segments were amplified for 10 min at 25°C, 60 min at 37°C and 25 min at 70°C. Samples were diluted 1 :10.
[0209] Table 15. Strand-specific primers for cDNA synthesis
[0210] Quantitative PCR
[0211] Quantitative real-time PCR was performed using a ViiA 7 Real-Time PCR System (Thermo Fisher Scientific). qPCR reactions (10 pl total) each contained LightCycler 480 SYBR Green I Master (Roche Diagnostics, #04887352001), 10 pM of each of the forward (strand specific) and reverse (ssTag) primers (Table 2), and 2 ng of diluted complementary DNA. Amplification of the 7SK region was used for normalization. PCR reactions used the following parameters: 1 cycle of 95°C for 5 min; 45 cycles of 95°C for 10 s, 58°C for 10 s and 72°C for 19 s; and a final melting curve cycle of 60°C to 97°C in 0.05°C / s steps. The quantification of each IGS was normalized for the 7SK region.
[0212] Table 16. Primers for qPCR
[0213] Example 14. Image Acquisition and dataset preparation for preleukemic cells.
[0214] Obtaining primary cell lines Hematopoietic Progenitor Stem Cells (HPSC) were extracted from female and male 8 -weeks-old mice and preleukemia induced through retroviral infection of a plasmid containing the MLL-AF9 translocation and respective empty vector (EV) as control. Preleukemic cells were then cultured for a maximum of 4 weeks and collected at different time points (day 0, 1 , 3.5, 7, 14, 21 , 28). Cells were then fixed with 4% PFA for 10 min on 8-well-chamber slides for imaging.
[0215] Datasets xvi, xvii and xviii
[0216] Dataset xvi consists of RNA Pol II phSer5 super resolution images of 370 preleukemic cells and 151 EV (empty vector, controls). Preleukemic cells are obtained by extracting Hematopoietic Progenitor Stem Cells (HPSC) from female and male 8-weeks-old C57BL / 6 mice which were induced preleukemia by retroviral infection of a plasmid containing the MLL-AF9 translocation.
[0217] Controls were generated by retroviral infection of empty vector (EV) to female and male 8-weeks- old C57BL / 6 mice. Preleukemic cells were then cultured for a maximum of 4 weeks and collected at different time points (day 0, 1 , 3.5, 7, 14, 21 , 28). Cells were then fixed with 4% PFA for 10 min on 8-well-chamber slides for imaging.
[0218] Dataset xvii consists of RNA Pol II phSer5 super resolution images of 120 preleukemic cells. Preleukemic cells are obtained by extracting Hematopoietic Progenitor Stem Cells (HPSC) from Phf19-KO mouse which were induced preleukemia by retroviral infection of a plasmid containing the MLL-AF9 translocation. Images of the cells were generated in an identical manner to dataset xvi.
[0219] Dataset xviii consists of DAPI stained nucleic acid confocal microscopy images of 606 preleukemic cells and 268 EV (empty vector, controls). Preleukemic and control cells and images are generated in an identical way to the dataset xvi.
[0220] Immunolabelling and super resolution
[0221] Cells were plated in chambered cover glass (p-Slide 8 Well Chamber Slide, Ibidi, #80826) at a concentration of 100,000 cells / cm2 in growth medium for 8-16h. Cells were fixed with 4% PFA (Leagene, #DF10315) for 10 min at room temperature and then rinsed with PBS. Cells were permeabilized with 0.4% (v / v) Triton X-100 (Sigma-Aldrich, #9036-19-5) in PBS for 15 min and then blocked in blocking buffer [10% bovine serum albumin [BSA], #B2064-50G; 0.01% (v / v) Triton X-100 in PBS] for 1 h at room temperature. Cells were incubated with primary antibodies [anti- RNA polymerase II CTD repeat YSPTSPS (phospho S5; Abeam #ab5408] in blocking buffer at 1 :50 dilution for STORM overnight at 4°C. On the second day, cells were washed three times for 5 min each with wash buffer (2% BSA and 0.01% Triton X-100 in PBS) and incubated in secondary antibody [goat anti-mouse IgG Alexa Fluor™ 647, ThermoScientific #A21235], added at 1 :250 dilution in blocking buffer for 1 h at room temperature. Cells were then washed three times for 5 min each with wash buffer, and once with PBS. Samples were stored in PBS at 4°C prior to imaging. For confocal imaging, DAPI was added after permeabilization and blocking for 10 min at room temperature, followed by 1 PBS wash of 5 min.
[0222] STORM imaging
[0223] STORM imaging was performed using a N-STORM 4.0 microscope (Nikon) with an iXon Ultra 897 camera (Andor) and a CFI HP Apochromat TIRF *100 1.49 oil objective, under highly inclined and laminated optical sheet (HILO) illumination mode. NIS element (4.60 and 5.21) software was used for image acquisition.
[0224] Continuous imaging acquisition was performed with simultaneous 405 and 647 nm illumination of the sample, 10 ms exposure time for 60,000 frames. A 647 nm laser was used at constant 3.5 kW / cm2, and 405 nm laser power was gradually increased over the imaging period.
[0225] During the super resolution imaging, an imaging buffer was added to the sample: 100 mM cysteamine MEA (Sigma-Aldrich, #30070), 1 % Glox solution (0.5 mg / ml glucose oxidase, 40 mg / ml catalase; Sigma-Aldrich, #G2133 and #C100), 5% glucose (Sigma-Aldrich, #G8270) in PBS.
[0226] N-STORM microscope has a resolution of 160 nm / pixel. Super-resolution images were analyzed and rendered in I nsight323 (gift of Bo Huang, UCSF). Specifically, point spread functions (PSFs) from the emission from single fluorophores were identified at individual frames of the acquired videos, based on a set threshold of intensity. They were fit to a two-dimensional Gaussian curve. From it, the centroid of the PSF was obtained, and the x and y positions for 2D STORM imaging were set with a precision of 9 decimals.
[0227] Confocal imaging
[0228] Confocal images of DNA by DAPI staining were performed with an LSM-980 Airyscan 2 inverted microscope, with an objective lens 63x / 1.4 Oil DIC, imaging software Zen. Pixel size 104nm / pixel.
[0229] Image Rendering
[0230] The molecule localization coordinates of the selected region of interest (ROI, i.e. nucleus) were multiplied by 10 corresponding to a resolution of 16 nm / px and then rendered into images by computing a Gaussian kernel density estimate using Scikit-learn KDTree functionl . To enable the model to differentiate between the background of the image outside the nucleus and regions within the nucleus lacking localizations, we set the background of full-size 10x rendered images to white. To avoid changing the aspect ratio, which would have changed the spatial relationship of the imaged molecules, each image was first resized to the desired dimension on its longest axis by keeping the aspect ratio and was subsequently padded with background color on the shortest axis.
[0231] Example 15. Retraining DenseNet-121 with preleukemic cell images
[0232] Dense121 architecture was retrained without reusing any previous weights. For the retraining phase, to ensure a balanced representation and robust model evaluation, dataset xvi, xvii and xviii were randomly partitioned such that images from each day (Figure 8) are included across all subsets — allocating 80% for the training set, 20% of the remaining data for validation, and then 20% of the subsequent remainder for the test set. By distributing images from each day across all dataset subsets, the model ensures to recognize and classify the conditions consistently, regardless of the day the images were captured. This approach helps to mitigate any biases that might arise from variations in experimental conditions or imaging across different days.
[0233] To assess the model's adaptability and effectiveness on different genetic backgrounds, the training set of dataset xvi was augmented with 6 images from dataset xvii. The model was then tested on the remaining images from dataset xvii to evaluate its predictive performance and robustness in identifying preleukemic states under altered genetic conditions. This approach allowed us to directly measure how well the model generalized across different experimental conditions and genetic variations.
[0234] Example 16. AINU trained with preleukemic cells identify preleukemic cells in a robust manner.
[0235] AINU model demonstrates robust performance in classifying preleukemic cells across RNA Pol II phSer5 super resolution images (datasets xvi, xvii) and DAPI stained DNA confocal images (dataset xviii). Specifically, when applied to dataset xvi, the model accurately classified 91 % of the cases achieving a ROC AUC (Receiver Operating Characteristic Area Under the Curve) of 0.91 (see Figure 9A, Figure 9B). The high level of accuracy highlights the model’s capability to identify early markers of leukemia with high precision.
[0236] Further testing on the training set additionally including 6 images from dataset xvii, which included images of preleukemic cells with an additional mutation, revealed that the addition of these images into the training set enables the AINU model to adapt to new pathological features effectively. This addition improved the model’s ability to generalize to unseen images of the mutated disease, maintaining a ROC AUC greater than 0.90 (Figure 9C, Figure 9D).
[0237] Moreover, the AINU model excelled in analyzing DNA images acquired through confocal microscopy. In this application, the model achieved an impressive ROC AUC of 0.97 and successfully recognized 98% of preleukemic cells (Figure 9E, Figure 9F).
[0238] These findings highlight the versatility of the AINU model and effectiveness across different types of biomedical imaging, establishing it as a powerful tool in the ongoing efforts to improve diagnostic accuracy in leukemia research.
[0239] Example 17. Preleukemic cells demonstrate denser RNA Pol II clusters
[0240] Cluster analysis involves examining localizations of the fluorescent molecules within a image to understand how molecules or particles are grouped or dispersed at a scale finer than what traditional light microscopy could resolve. To further investigate the biological signature and to better understand the model decision, cluster analysis and other statistical analysis were performed (Figure 10). As shown in Figure 10, density of localizations (normalized per area or nucleus) increases in preleukemic compared to normal cells, while the NND (nearest neighbour distance) gets smaller between clusters of preleukemic cells, suggesting denser and more compacted areas of active RNA Polymerase II in preleukemia (Figure 11).
[0241] Centroid of a Cluster
[0242] Cluster analysis was conducted following the methods outlined in Ricci et al., 2015. Specifically, the localization lists were binned to construct discrete localization images with a pixel size of 10 nm and pixel intensity equal to the number of localizations falling in the corresponding pixel. These images were convoluted with a 5x5 pixels kernel to obtain density maps and transformed into binary images by applying a constant threshold, such that each pixel has a value of either 1 if the density surpasses the threshold value and 0 if not. Localizations falling inside a 0-value area of the image were discarded. Cluster centroids were obtained from local maxima of the density map, and localizations were assigned to clusters using a distance-based algorithm. New cluster centroids were then calculated as the average of the localizations previously assigned to the cluster. This process was repeated iteratively until convergence of the sum of the squared distances of the localizations belonging to the cluster and the centroid. The final centroid position and number of localizations were obtained for each cluster. Cluster sizes were calculated as the standard deviation of x-y coordinates from the centroid of the cluster.
[0243] Centroid of a Cluster
[0244] The centroid of a cluster is calculated as the mean of the x and y coordinates of all the localizations within that cluster. This serves as the geometric center of the cluster and is crucial for further spatial analyses, including calculations of inter-cluster distances.
[0245] Area of a Cluster (nm2)
[0246] The area of a cluster is derived from the spread of its localizations. Assuming the cluster forms roughly a circular shape, the area (A) can be estimated using the formula:
[0247] A = nr2where r is the radius of the cluster. The radius is approximated as the average of the standard deviations of the x and y coordinates of the localizations, expressed as:
[0248] Density of a cluster The density of a cluster provides insights into how tightly localizations are packed within the cluster area. It is calculated as:
[0249] Number of localizations Density = - - - where A is the area of the cluster as defined above. where (Xa, Ta) and (Xb, Yb~) are the coordinates of the centroids of two neighboring clusters.
[0250] Nearest Neighbour Distance (NND) between clusters’ centroids was calculated by knnsearch.m MATLAB function and the NND histogram of experimental data was obtained by considering all the NNDs of individual nuclei. Plots and statistical analysis were performed on R using the Mann-Witney test.
[0251] Example 18. Interpretable Al highlights regions of high RNA Pol II intensity in the nucleus of preleukemic cells.
[0252] To evaluate the contribution of each input feature (image pixels) on the model’s output and prediction accuracy, and to gain an understanding of the model's decisions, Gradient-weighted Class Activation Mapping (Grad-CAM), a form of interpretable Al, was employed.
[0253] Grad-CAM provides a visual explanation for the decisions made by CNN by generating a heatmap overlay on the input image (Figure 12). Grad-CAM was conducted following the methods outlined in Selvaraju et al., 2017. The heatmap generated from Grad-CAM highlights the regions within the image that were most influential in the model's predictions, and verifies whether the model focuses on relevant features for classification, thereby aiding in the validation and refinement of the model's performance.
[0254] The heatmap highlights spots of high intensity of RNA Polymerase II in the nucleus of preleukemic cells (Figure 12). The highlighted area also emphasise regions that significantly impact the model’s decision. However, further investigation is needed to explain the relationship between the spatial organization of molecules and their impact on biological functions.
[0255] CLAUSES:
[0256] 1. A method for identifying cellular heterogeneity within a population of cells, the method comprising: (a) acquiring a plurality of images of one or more cells;
[0257] (b) defining a test set comprising one or more of the acquired images;
[0258] (c) providing the test set to a trained artificial neural network;
[0259] (d) applying the trained artificial neural network to classify one or more cells in one or more images of the test set; and
[0260] (e) identifying one or more features associated with cellular heterogeneity in one or more classified cells of the one or more images of the test set. la. The method of Clause 1 , wherein the artificial neural network is trained by:
[0261] (b) defining a training set comprising of one or more of the acquired images, defining a validation set comprising one or more of the acquired images that are not within the training set and optionally defining a test set comprising one or more of the acquired images that are not within the training set or the validation set;
[0262] (c1) initialising weights associated with a plurality of filters of the artificial neural network;
[0263] (c2) applying the plurality of filters to the training set and defining the weight of each of the plurality of filters;
[0264] (c3) applying the plurality of filters having the defined weights to the validation set and assigning a classification and predicted value of one or more cells in the images of the validation set;
[0265] (c4) applying a cost function to determine a loss value, wherein the loss value is defined as a difference between the predicted value of the assigned classification and the true value for the one or more cells in the images of the validation set from step (c3);
[0266] (c5) updating the weight of each of the plurality of filters of the artificial neural network;
[0267] (c6) re-applying the plurality of filters to the validation set and re-assigning a classification and new predicted value of the one or more cells in the images of the validation set;
[0268] (c7) re-applying the cost function to obtain a reduced loss value defined as a difference between the new predicted value from step (c6) and the true value for the one or more cells in the images of the validation set. l b. The method of Clause 1 a, wherein the loss function is Cross Entropy Loss, Mean Squared Error Loss, or Binary Cross Entropy Loss.
[0269] 2. The method of Clause 1a or Clause 1 b, wherein a filter performs an operation comprising a matrix of weights that is systematically applied to one or more images or a layer of the artificial neural network.
[0270] 3. The method of Clause 2, wherein the operation is a convolution.
[0271] 4. The method of any of Clauses 1 a to 3, wherein initialising weights comprises assigning random values for each weight. 4a. The method of any of Clauses 1a to 3, wherein initialising weights comprises Xavier initialization.
[0272] 4b. The method of any of Clauses 1a to 3, wherein initialising weights is based on a pre-trained artificial neural network.
[0273] 5. The method of any of Clauses 1a to 3, wherein initialising weights comprises assigning set values for at least one weight of a filter.
[0274] 6. The method of any of Clauses 1a to 5, wherein reducing the loss value involves back- propagating the difference calculated at step (c4) into the neural network.
[0275] 7. The method of any of Clauses 1a, wherein the reduced loss value from (c7) reaches a minimised loss value, wherein the minimised loss value is obtained from re-applying the cost function.
[0276] 8. The method of any of Clauses 1a to 7, wherein the reduced loss value is obtained by a process of gradient descent.
[0277] 9. The method of any of Clauses 1 a to 8, wherein updating the weights of step (c5) comprises assigning new weights to the plurality of filters.
[0278] 10. The method of any of Clauses 1a to 9, wherein steps (c5) to (c7) are repeated until the updated weights of the plurality of filters do not improve the classification beyond a predefined threshold.
[0279] 11 . The method of any of Clauses 1 a to 9, wherein steps (c5) to (c7) are repeated until the desired number of epochs is reached.
[0280] 12. The method of any of Clauses 1 a to 9, wherein steps (c5) to (c7) are repeated until a minimum loss value is reached.
[0281] 13. A method for training an artificial neural network for identifying cellular heterogeneity within a population of cells, the method comprising:
[0282] (a) acquiring a plurality of images of one or more cells;
[0283] (b) defining a training set comprising of one or more of the acquired images, defining a validation set comprising one or more of the acquired images that are not within the training set and optionally defining a test set comprising one or more of the acquired images that are not within the training set or the validation set;
[0284] (c) providing the training set to an artificial neural network to identify cellular heterogeneity in one or more cells of the training set;
[0285] (d) optionally applying the trained artificial neural network to classify one or more cells of the images of the test set; and (e) optionally identifying one or more features associated with cellular heterogeneity in one or more classified cells of the images of the test set.
[0286] 13a. The method of Clause 13, wherein training the artificial neural network comprises:
[0287] (c1) initialising weights associated with a plurality of filters of the artificial neural network;
[0288] (c2) applying the plurality of filters to the training set and defining the weight of each of the plurality of filters;
[0289] (c3) applying the plurality of filters having the defined weights to the validation set and assigning a classification and predicted value of one or more cells in the images of the validation set;
[0290] (c4) applying a cost function to determine a loss value, wherein the loss value is defined as a difference between the predicted value of the assigned classification and the true value for the one or more cells in the images of the validation set from step (c3);
[0291] (c5) updating the weight of each of the plurality of filters of the artificial neural network;
[0292] (c6) re-applying the plurality of filters to the validation set and re-assigning a classification and new predicted value of the one or more cells in the images of the validation set;
[0293] (c7) re-applying the cost function to obtain a reduced loss value defined as a difference between the new predicted value from step (c6) and the true value for the one or more cells in the images of the validation set;
[0294] 14. A method for identifying cellular heterogeneity within a population of cells, the method comprising:
[0295] (a) acquiring a plurality of images of one or more cells;
[0296] (b) applying a trained artificial neural network to classify one or more cells of the plurality of images;
[0297] (c) identifying one or more features associated with cellular heterogeneity in one or more classified cells of the plurality of images.
[0298] 14a. The method of Clause 14, wherein the artificial neural network is trained according to the method of Clause 13a.
[0299] 15. A method for determining the progression of a cancer within a population of cells, the method comprising:
[0300] (a) acquiring a plurality of images of one or more cells;
[0301] (b) defining a test set comprising one or more of the acquired images;
[0302] (c) applying a trained artificial neural network to classify one or more cells in one or more images of the test set; and
[0303] (d) identifying one or more features associated with cellular heterogeneity in a cancer cell to determine the progression of the cancer in one or more classified cells of the images of the test set. 15a. The method of Clause 15, wherein step (b) comprises:
[0304] (b1) defining a training set comprising of one or more of the acquired images, defining a validation set comprising one or more of the acquired images that are not within the training set and defining a test set comprising one or more of the acquired images that are not within the training set or the validation set.
[0305] 15b. The method of Clause 15a, further comprising training the artificial neural network prior to step (c), by:
[0306] (c1) initialising weights associated with a plurality of filters of the artificial neural network;
[0307] (c2) applying the plurality of filters to the training set and defining the weight of each of the plurality of filters;
[0308] (c3) applying the plurality of filters having the defined weights to the validation set and assigning a classification and predicted value of one or more cells in the images of the validation set;
[0309] (c4) applying a cost function to determine a loss value, wherein the loss value is defined as a difference between the predicted value of the assigned classification and the true value for the one or more cells in the images of the validation set from step (c3);
[0310] (c5) updating the weight of each of the plurality of filters of the artificial neural network;
[0311] (c6) re-applying the plurality of filters to the validation set and re-assigning a classification and new predicted value of the one or more cells in the images of the validation set;
[0312] (c7) re-applying the cost function to obtain a reduced loss value defined as a difference between the new predicted value from step (c6) and the true value for the one or more cells in the images of the validation set.
[0313] 15c. A method for detecting preleukemic cells within a population of cells, the method comprising:
[0314] (a) acquiring a plurality of images of one or more cells;
[0315] (b) defining a test set comprising one or more of the acquired images;
[0316] (c) applying a trained artificial neural network to classify one or more cells in one or more images of the test set; and
[0317] (d) identifying one or more features associated with cellular heterogeneity in a leulemic cell to detect preleukemic cells in one or more classified cells of the images of the test set.
[0318] 15d. The method of Clause 15c, wherein step (b) comprises:
[0319] (b1) defining a training set comprising of one or more of the acquired images, defining a validation set comprising one or more of the acquired images that are not within the training set and defining a test set comprising one or more of the acquired images that are not within the training set or the validation set.
[0320] 15e. The method of Clause 15c or Clause 15d, further comprising training the artificial neural network prior to step (c), by:
[0321] (c1) initialising weights associated with a plurality of filters of the artificial neural network; (c2) applying the plurality of filters to the training set and defining the weight of each of the plurality of filters;
[0322] (c3) applying the plurality of filters having the defined weights to the validation set and assigning a classification and predicted value of one or more cells in the images of the validation set;
[0323] (c4) applying a cost function to determine a loss value, wherein the loss value is defined as a difference between the predicted value of the assigned classification and the true value for the one or more cells in the images of the validation set from step (c3);
[0324] (c5) updating the weight of each of the plurality of filters of the artificial neural network;
[0325] (c6) re-applying the plurality of filters to the validation set and re-assigning a classification and new predicted value of the one or more cells in the images of the validation set;
[0326] (c7) re-applying the cost function to obtain a reduced loss value defined as a difference between the new predicted value from step (c6) and the true value for the one or more cells in the images of the validation set.
[0327] 16. A method for identifying the progression of cancer within a population of cells, comprising:
[0328] (a) acquiring a plurality of images of one or more cells of the population of cells;
[0329] (b) applying a trained artificial neural network to classify one or more cells of the plurality of images;
[0330] (c) identifying one or more features associated with the progression of cancer in one or more classified cells of the plurality of images.
[0331] 16a. The method of Clause 16, wherein the trained artificial neural network is trained by:
[0332] (a1) defining a training set comprising of one or more of the acquired images, defining a validation set comprising one or more of the acquired images that are not within the training set; and
[0333] (a2) training an artificial neural network by:
[0334] (a3) initialising weights associated with a plurality of filters of the artificial neural network;
[0335] (a4) applying the plurality of filters to the training set and defining the weight of each of the plurality of filters;
[0336] (a5) applying the plurality of filters having the defined weights to the validation set and assigning a classification and predicted value of one or more cells in the images of the validation set;
[0337] (a6) applying a cost function to determine a loss value, wherein the loss value is defined as a difference between the predicted value of the assigned classification and the true value for the one or more cells in the images of the validation set from step (c3);
[0338] (a7) updating the weight of each of the plurality of filters of the artificial neural network;
[0339] (a8) re-applying the plurality of filters to the validation set and re-assigning a classification and new predicted value of the one or more cells in the images of the validation set; and (a9) re-applying the cost function to obtain a reduced loss value defined as a difference between the new predicted value from step (c6) and the true value for the one or more cells in the images of the validation set; and
[0340] (a 10) optionally identifying one or more features associated with cancer progression in one or more classified cells of the images of the test set.
[0341] 17. A method for diagnosing a disease within a population of cells, comprising:
[0342] (a) acquiring a plurality of images of one or more cells;
[0343] (b) defining a test set comprising one or more of the acquired images;
[0344] (c) applying a trained artificial neural network to classify one or more cells in one or more images of the test set; and
[0345] (d) identifying one or more features associated with cellular heterogeneity to diagnose the disease in one or more classified cells of the images of the test set.
[0346] 17a. The method of Clause 17, wherein step (b) comprises:
[0347] (b1) defining a training set comprising of one or more of the acquired images, defining a validation set comprising one or more of the acquired images that are not within the training set and defining a test set comprising one or more of the acquired images that are not within the training set or the validation set.
[0348] 17b. The method of Clause 17a, further comprising training the artificial neural network prior to step (c), by:
[0349] (c1) initialising weights associated with a plurality of filters of the artificial neural network;
[0350] (c2) applying the plurality of filters to the training set and defining the weight of each of the plurality of filters;
[0351] (c3) applying the plurality of filters having the defined weights to the validation set and assigning a classification and predicted value of one or more cells in the images of the validation set;
[0352] (c4) applying a cost function to determine a loss value, wherein the loss value is defined as a difference between the predicted value of the assigned classification and the true value for the one or more cells in the images of the validation set from step (c3);
[0353] (c5) updating the weight of each of the plurality of filters of the artificial neural network;
[0354] (c6) re-applying the plurality of filters to the validation set and re-assigning a classification and new predicted value of the one or more cells in the images of the validation set;
[0355] (c7) re-applying the cost function to obtain a reduced loss value defined as a difference between the new predicted value from step (c6) and the true value for the one or more cells in the images of the validation set.
[0356] 18. A method for diagnosing a disease within a population of cells, comprising:
[0357] (a) acquiring a plurality of images of one or more cells; (b) applying the trained artificial neural network of any of Claims 1 a to 12 to classify one or more cells of the plurality of images;
[0358] (c) identifying one or more features associated with the diagnosis of disease in one or more classified cells of the plurality of images.
[0359] 19. A method for determining an infection within a population of cells, comprising:
[0360] (a) acquiring a plurality of images of one or more cells;
[0361] (b) defining a test set comprising one or more of the acquired images;
[0362] (c) providing the test set to a trained artificial neural network;
[0363] (d) applying the trained artificial neural network to classify one or more cells in one or more images of the test set;
[0364] (e) identifying one or more features associated with cellular heterogeneity in an infected cell to determine the presence of infection in one or more classified cells of the images of the test set.
[0365] 20. A method for determining an infection within a population of cells, comprising:
[0366] (a) acquiring a plurality of images of one or more cells;
[0367] (b) applying a trained artificial neural network to classify one or more cells of the plurality of images;
[0368] (c) identifying one or more features associated with the determining the infection in one or more classified cells of the plurality of images.
[0369] 20a. The method of Clause 20, wherein between step (a) and step (b) the method comprises:
[0370] (a1) defining a training set comprising of one or more of the acquired images, defining a validation set comprising one or more of the acquired images that are not within the training set and defining a test set comprising one or more of the acquired images that are not within the training set or the validation set.
[0371] 20b. The method of Clause 20a, further comprising training the artificial neural network prior to step (b), by:
[0372] (a2) initialising weights associated with a plurality of filters of the artificial neural network;
[0373] (a3) applying the plurality of filters to the training set and defining the weight of each of the plurality of filters;
[0374] (a4) applying the plurality of filters having the defined weights to the validation set and assigning a classification and predicted value of one or more cells in the images of the validation set;
[0375] (a5) applying a cost function to determine a loss value, wherein the loss value is defined as a difference between the predicted value of the assigned classification and the true value for the one or more cells in the images of the validation set from step (c3);
[0376] (a6) updating the weight of each of the plurality of filters of the artificial neural network; (a7) re-applying the plurality of filters to the validation set and re-assigning a classification and new predicted value of the one or more cells in the images of the validation set;
[0377] (a8) re-applying the cost function to obtain a reduced loss value defined as a difference between the new predicted value from step (c6) and the true value for the one or more cells in the images of the validation set.
[0378] 21 . The method of any of Clauses 1 to 20, wherein the population of one or more cells comprises, consists essentially of or consists of a single cell type.
[0379] 22. The method of any of Clauses 1 to 20, wherein the population of one or more cells comprises multiple cell types.
[0380] 23. The method of any of Clauses 1 to 22, wherein identifying cellular heterogeneity involves identifying one or more different cell types.
[0381] 24. The method of any of Clauses 1 to 23, wherein identifying cellular heterogeneity involves identifying differentiation states of one or more cells.
[0382] 25. The method of any of Clauses 1 to 24, wherein identifying cellular heterogeneity involves identifying degrees of chromatin organisation.
[0383] 26. The method of any of Clauses 1 to 24, wherein identifying cellular heterogeneity involves identifying subcellular / cellular features.
[0384] 26a. The method of any of Clauses 1 to 24, wherein identifying cellular heterogeneity involves identifying one or more features used to identify different cell types.
[0385] 27. The method of any of Clauses 1 to 24, wherein identifying cellular heterogeneity involves identifying one or more features responsive to drug treatment.
[0386] 28. The method of any of Clauses 1 to 24, wherein identifying cellular heterogeneity involves identifying one or more features indictive of an extent of infection or disease.
[0387] 29. The method of any of Clauses 1 to 24, wherein identifying cellular heterogeneity involves identifying one or more features associated with cancer cell pluripotency.
[0388] 29a. The method of any of Clauses 1 to 24, wherein identifying cellular heterogeneity involves identifying one or more features indicative of presence of premalignant cells. 29b. The method of any of Clauses 1 to 24, wherein identifying cellular heterogeneity involves identifying one or more features indicative of presence of leukemia or premalignant leukemia cells; particularly premalignant acute myeloid leukemia cells.
[0389] 29c. The method of any of Clauses 1 to 29b, wherein identifying cellular heterogeneity involves combination of any of the features of Clauses 25 to 29b.
[0390] 30. The method of any of Clauses 1 to 29, wherein the acquired images are super resolution images.
[0391] 31. The method of Clause 30, wherein the super resolution images are acquired by Singlemolecule localization microscopy (SMLM), Stochastic Optical Reconstruction Microscopy (STORM), Photoactivated Localization Microscopy (PALM), DNA points accumulation for imaging in nanoscale topography (DNA-PAINT) or a Stimulated emission depletion (STED) microscopy.
[0392] 31 a. The method of Clause 31 , wherein the resolution of the STORM imaging is at least 20 nm in a lateral direction and at least and 60 nm in an axial direction.
[0393] 32. The method of any of Clauses 1 to 31 a, further comprising: generating a further plurality of images of one or more cells based on one or more of the acquired images.
[0394] 33. The method of Clause 32, wherein the further plurality of images are not super resolution images.
[0395] 33a. The method of Clause 32 or Clause 33, wherein the further plurality of images are obtained from diffraction limited microscopy, such as confocal microscopy images or widefield microscopy images.
[0396] 34. The method of Clause 32, wherein the further plurality of images are super resolution images.
[0397] 35. The method of Clause 32, wherein the further plurality of images are simulated images.
[0398] 35a. The method of Clause 35, wherein the simulated images are generated from one or more randomly selected image of the plurality of images of step (a).
[0399] 36. The method of any of Clauses 32 to 35a, wherein the artificial neural network is trained using a training set comprising one or more of the further plurality of images.
[0400] 37. The method of Clause 35, Clause 35a or Clause 36 when dependent on Clause 35 or Clause 35a, wherein a generative Al approach is used to generate the simulated images. 38. The method of Clause 37, wherein the generative Al is a generative adversarial network (GAN) or a Variational autoencoder (VAE).
[0401] 39. The method of Clause 38, wherein the trained artificial neural network is trained using simulated images based on randomly selected images of the plurality of images of step (a).
[0402] 40. The method of any of Clauses 32 to 39, wherein the number of generated images is more than about 1000, more than about 3000, more than about 5000, more than about 7000, more than about 9000, or more than about 11000.
[0403] 41 . The method of any of Clauses 32 to 40, wherein the number of generated images is less than about 20000, less than about 18000, less than about 16000, or less than about 14000.
[0404] 42. The method of any of Clauses 32 to 41 , wherein the number of generated images is between about 1000 and 20000, between about 3000 and 18000, between about 5000 and 16000; or between about 7000 and 14000.
[0405] 43. The method of any of Clauses 1 to 42, wherein one or more molecules of interest in the population of cells are labelled prior to image acquisition.
[0406] 44. The method of Clause 43, wherein labelling involves immunolabelling.
[0407] 45. The method of Clause 43, wherein labelling involves click chemistry.
[0408] 45a. The method of Clause 43, wherein labelling involves DAPI staining.
[0409] 46. The method of Clause 43, wherein labelling involves binding of one ora plurality of fluorophores to the one or more molecules of interest.
[0410] 47. The method of Clause 46, wherein labelling involves binding of a plurality of fluorophores wherein different ones of the fluorophores have different emission wavelengths and wherein the different ones of the fluorophores are bound to different molecules of interest.
[0411] 48. The method of any of Clauses 1 a to 13a, 15a, 15b, 16a, 17a, 17b, 20a, 20b or any of Clauses 21 to 47 when dependent on any one of Clauses 1a to 13a, 15a, 15b, 16a, 17a, 17b, 20a, or 20b, wherein the number of images in the training set is at least about 50, at least about 75, at least about 100, at least about 150, at least about 250, at least about 500, at least about 750 or at least about 1000. 49. The method of any of Clauses 1a to 13a, 15a, 15b, 16a, 17a, 17b, 20a, 20b or any of Clauses 21 to 48 when dependent on any one of Clauses 1a to 13a, 15a, 15b, 16a, 17a, 17b, 20a, or 20b, wherein the number of images in the training set is less than about 2000, less than about 1500, less than about 1250, less than about 1000, or less than about 750.
[0412] 50. The method of any of Clauses 32 to 42, or any of Clauses 43 to 49 when dependent on any of Clauses 32 to 42, wherein the number of further images generated is less than about 1500, less than about 1000, less than about 750, less than about 500 or less than about 250.
[0413] 51 . The method of any of Clauses 32 to 42, or any of Clauses 43 to 50 when dependent on any of Clauses 32 to 42, wherein the number of further images generated is at least about 100, at least about 200, at least about 300, at least about 400, or at least about 500.
[0414] 52. The method of any of Clauses 32 to 42, or any of Clauses 43 to 51 when dependent on any of Clauses 32 to 42, wherein the number of further images generates is between about 100 and 1500, between about 200 and 1000, between about 300 and 750, or between about 400 and 500.
[0415] 53. The method of any of Clauses 1 to 52, wherein the training set includes up to about 90%, up to about 80%, up to about 70%, up to about 60% or up to about 50% of the acquired plurality of images.
[0416] 54. The method of any of Clauses 1 to 53, wherein the training set includes at least about 10%, at least about 20%, at least about 30%, at least about 40% or at least about 50% of the acquired plurality of images.
[0417] 55. The method of any of Clauses 1 to 54, wherein the training set includes between about 10% and 90%, between about 20% and 80%, between about 30% 70%, or between about 40% and 60% of the acquired plurality of images.
[0418] 56. The method of any of Clauses 1 to 55, wherein the test set includes up to about up to about 50%, up to about 40%, up to about 30%, up to about 20%, or up to about 10% of the acquired plurality of images; and / or at least about 5%, at least about 10%, at least about 15%, or at least about 20% of the acquired plurality of images; and / or between about 5% and 50%, between about 10% and 40% or between about 15% and 30% of the acquired plurality of images.
[0419] 57. The method of Clauses 1 to 56, wherein the images of the training set are selected by random sampling of the acquired plurality of images.
[0420] 58. The method of Clauses 1 to 57, wherein the training set includes at least one lower magnification image that frames one or more cells, and at least one higher magnification image that enlarges a part of one or more cells. 59. The method of Clauses 1 to 58, wherein supervised machine learning is used to identify a feature associated with cellular heterogeneity in one or more classified cells from one or more images of the training set.
[0421] 60. The method of any of Clauses 43 to 47 or any of Clauses 48 to 59 when dependent on any of Clauses 43 to 47, wherein the feature or features are based on the distribution of individual labels in one or more cells.
[0422] 61 . The method of Clause 60, wherein the distribution of individual labels is expressed in a graph network or a point cloud function.
[0423] 62. The method of Clause 60 or Clause 61 , wherein the distribution of individual labels indicates qualities such as density or length or shape of the labelled molecule of interest.
[0424] 63. The method of any of Clauses 60 to 62, wherein the labelled molecule of interest is core histone H3, RNA polymerase II and / or DNA; optionally wherein the feature is chromatin structure.
[0425] 64. The method of any of Clauses 1 to 63, wherein the feature relates to nucleosomal clutches, wherein a nucleosomal clutch is an arrangement or grouping of nucleosomes.
[0426] 65. The method of any of Clauses 1 to 64, comprising selecting an artificial neural network to identify cellular heterogeneity within a population of cells based on evaluation metrics.
[0427] 66. The method of Clause 65, wherein an artificial neural network is selected from known artificial neural network model architectures.
[0428] 67. The method of Clause 65 or Clause 66, wherein selecting an artificial neural network involves training a plurality of artificial neural network model architectures with a training set comprising of a one or more of the acquired images and selecting one of the trained artificial neural network model architectures.
[0429] 68. The method of any of Clauses 65 to 67 when dependent on any of Clauses 43 to 47 or 58 to 64, wherein selecting an artificial neural network involves training known artificial neural network model architectures using a distribution of individual labels from the training set.
[0430] 69. The method of any of Clauses 65 to 68, wherein training different artificial neural network model architectures involves applying an identical set of hyperparameters to each of the artificial neural network model architectures and performing K-fold cross validation. 70. The method of Clause 69, wherein the hyperparameters include epochs, batch size, learning rate and a type of optimiser.
[0431] 71 . The method of Clause 70, wherein the type of optimiser comprises one or more of: stochastic gradient descent, ADAM, ADAMW, Adadelta, and RMSprop.
[0432] 72. The method of Clause 70 or Clause 71 , wherein the learning rate is at least about 0.0001 , at least about 0.0005, at least about 0.001 or at least about 0.005.
[0433] 73. The method of any of Clauses 70 to 72, wherein the learning rate is less than about 0.1 , less than about 0.05, or less than about 0.01 .
[0434] 74. The method of any of Clauses 70 to 73, wherein the learning rate is between about 0.0001 and 0.1 , between about 0.0005 and 0,05, or between about 0.001 and 0.01.
[0435] 75. The method of any of Clauses 65 to 74, wherein the evaluation metrics comprise one or more of: F1 score, validation accuracy and loss, area under the Receiver Operating Characteristic curve, precision and recall.
[0436] 76. The method of any of Clauses 69 to 74, or Clause 75 when dependent on any of Clauses 69 to 74, wherein the selecting an artificial neural network comprises selecting a model with the highest average validation accuracy and smallest average loss area under the Receiver Operating Characteristic curve from K-fold cross validation.
[0437] 77. The method of any of Clauses 69 to 74, 76, or Clause 75 when dependent on any of Clauses 69 to 74, wherein hyperparameters of different values are tested on the selected artificial neural network model architecture.
[0438] 78. The method of any of Clauses 65 to 77, wherein training different artificial neural network model architectures involves applying at least one data augmentation technique.
[0439] 79. The method of Clause 78, wherein a data augmentation technique is applied to increase the size of the dataset and prevent overfitting.
[0440] 80. The method of any of Clauses 78 and 79, wherein the data augmentation technique involves applying Albumentations library, which comprises one or more of: horizontal I vertical flipping, random rotation within a range of 180 degrees, random changes in hue, saturation and values within 20% of the actual values and regions dropout.
[0441] 81. The method of any of Clauses 65 to 80, wherein the selected artificial neural network is retrained with a test set comprising one or more of the acquired images. 82. The method of any of Clauses 1 to 81 , wherein applying the trained artificial neural network to classify one or more cells in one or more images is based on a classification threshold, wherein the classification threshold is defined as a minimum required proportion of the images that are correctly classified using the trained artificial neural network.
[0442] 83. The method of Clause 82, wherein a classification threshold is determined by optimising true positive rate and false positive rate from a receiver operating characteristic (ROC) curve.
[0443] 84. The method of Clause 82 or Clause 83, wherein the classification threshold is at least about 0.5, at least about 0.6. at least about 0.7, at least about 0.8 or at least about 0.9.
[0444] 85. The method of any of Clauses 82 to 84, wherein the classification threshold is at least about 0.5, at least about 0.55, at least about 0.6, or at least about 0.65.
[0445] 86. The method of any of Clauses 82 to 85, wherein the classification threshold is less than about 1 , less than about 0.9, less than about 0.8, less than about 0.7, or less than about 0.6.
[0446] 87. The method of any of Clauses 82 to 86, wherein the classification threshold is between about 0.5 and 1 , between about 0.5 and 0.9, between about 0.5 and 0.8, between about 0.5 and 0.7, or between about 0.5 and 0.6.
[0447] 88. The method of any of Clauses 82 to 87, wherein the classification threshold is defined as an area under the ROC curve.
[0448] 89. The method of Clause 88, wherein the area under the ROC curve is at least 0.7, at least about 0.8, or at least 0.9.
[0449] 90. The method of any of Clauses 1 to 89, wherein the feature or features are indicative of one or more cell types and / or pathologies.
[0450] 91 . The method of any of Clauses 1 to 90, wherein the feature or features are indicative of one or more differentiation states of a cell.
[0451] 92. The method of any of Clauses 1 to 91 , wherein the feature or features are indicative of chromatin organisation in a cell.
[0452] 93. The method of any of Clauses 1 to 92, wherein the feature or features are indicative of transcriptional activity in a cell. 94. The method of any of Clauses 1 to 93, wherein the feature or features are indicative of subcellular I cellular characteristics of a cell.
[0453] 95. The method of any of Clauses 1 to 94, wherein the feature or features are indicative of a type of cancer and / or cancer progression of a cell.
[0454] 95a. The method of any of Clauses 1 to 95, wherein the feature or features are indicative of a precancerous characteristic in a cell, such as a preleukemic characteristic of a cell.
[0455] 96. The method of any of Clauses 1 to 95, wherein the feature or features are indicative of pluripotency of a cell.
[0456] 97. The method of any of Clauses 1 to 96, wherein the feature or features are indicative of a disease state or an extent of disease progression of a cell.
[0457] 98. The method of any of Clauses 1 to 97, wherein the feature or features are indicative of a response to drug treatment of a cell.
[0458] 99. The method of any of Clauses 1 to 98, wherein the feature or features are indicative of the viral infection level of a cell.
[0459] 100. The method of any of Clauses 1 to 99, wherein identifying one or more features associated with cellular heterogeneity in one or more classified cells further comprises: identifying an area of one or more images used for classification by applying a primary attribution algorithm; and determining one or more features of interest from the identified area of the one or more images.
[0460] 101. The method of Clause 100, wherein the primary attribution algorithm assigns positive and negative attribution scores to areas of the one or more images.
[0461] 102. The method of Clause 100 or Clause 101 , wherein the feature or features of interest are determined from areas of the one or more images with positive attribution scores.
[0462] 103. The method of any of Clauses 100 to 102, wherein the primary attribution algorithm is an occlusion based algorithm.
[0463] 104. The method of any of Clauses 100 to 103, wherein the primary attribution algorithm is validated by a class activation mapping method. 105. The method of any of Clauses 1 to 104, wherein the feature is a biological feature, an endogenous feature of a cell, or an exogenous feature of a cell.
[0464] 106. A system for identifying cellular heterogeneity within a population of cells, comprising: a database library that includes a training set comprising of one or more of the acquired images, a validation set comprising one or more of the acquired images that are not within the training set and a test (or sample) set comprising one or more of the acquired images that are not within the training set or the validation set; a processor in communication with the database, the processor being adapted to: receive and provide the test set to a trained artificial neural network; apply the trained artificial neural network to classify one or more cells in one or more images of the test set; and identify one or more features associated with cellular heterogeneity in one or more classified cells of the one or more images of the test set.
[0465] 107. The system of Clause 106, wherein the artificial neural network is trained by: defining a training set comprising of one or more of the acquired images, defining a validation set comprising one or more of the acquired images that are not within the training set and optionally defining a test set comprising one or more of the acquired images that are not within the training set or the validation set; initialising weights associated with a plurality of filters of the artificial neural network; applying the plurality of filters to the training set and defining the weight of each of the plurality of filters; applying the plurality of filters having the defined weights to the validation set and assigning a classification and predicted value of one or more cells in the images of the validation set; applying a cost function to determine a loss value, wherein the loss value is defined as a difference between the predicted value of the assigned classification and the true value for the one or more cells in the images of the validation set from step (c3); updating the weight of each of the plurality of filters of the artificial neural network; re-applying the plurality of filters to the validation set and re-assigning a classification and new predicted value of the one or more cells in the images of the validation set; re-applying the cost function to obtain a reduced loss value defined as a difference between the new predicted value from step (c6) and the true value for the one or more cells in the images of the validation set.
Claims
CLAIMS1. A method for identifying cellular heterogeneity within a population of cells, the method comprising:(a) acquiring a plurality of images of one or more cells;(b) defining a test set comprising one or more of the acquired images;(c) providing the test set to a trained artificial neural network;(d) applying the trained artificial neural network to classify one or more cells in one or more images of the test set; and(e) identifying one or more features associated with cellular heterogeneity in one or more classified cells of the one or more images of the test set.
2. The method of Claim 1 , wherein the artificial neural network is trained by:(a) defining a training set comprising of one or more of the acquired images, defining a validation set comprising one or more of the acquired images that are not within the training set and optionally defining a test set comprising one or more of the acquired images that are not within the training set or the validation set;(b) initialising weights associated with a plurality of filters of the artificial neural network;(c) applying the plurality of filters to the training set and defining the weight of each of the plurality of filters;(d) applying the plurality of filters having the defined weights to the validation set and assigning a classification and predicted value of one or more cells in the images of the validation set;(e) applying a cost function to determine a loss value, wherein the loss value is defined as a difference between the predicted value of the assigned classification and the true value for the one or more cells in the images of the validation set from step (c3);(f) updating the weight of each of the plurality of filters of the artificial neural network;(g) re-applying the plurality of filters to the validation set and re-assigning a classification and new predicted value of the one or more cells in the images of the validation set;(h) re-applying the cost function to obtain a reduced loss value defined as a difference between the new predicted value from step (c6) and the true value for the one or more cells in the images of the validation set.
3. The method of any of Claim 2, wherein the reduced loss value from (h) reaches a minimised loss value, wherein the minimised loss value is obtained from re-applying the cost function.
4. A method for training an artificial neural network for identifying cellular heterogeneity within a population of cells, the method comprising:(a) acquiring a plurality of images of one or more cells;(b) defining a training set comprising of one or more of the acquired images, defining a validation set comprising one or more of the acquired images that are not within the training setand optionally defining a test set comprising one or more of the acquired images that are not within the training set or the validation set;(c) providing the training set to an artificial neural network to identify cellular heterogeneity in one or more cells of the training set;(d) optionally applying the trained artificial neural network to classify one or more cells of the images of the test set; and(e) optionally identifying one or more features associated with cellular heterogeneity in one or more classified cells of the images of the test set.
5. A method for identifying cellular heterogeneity within a population of cells, the method comprising:(a) acquiring a plurality of images of one or more cells;(b) applying a trained artificial neural network to classify one or more cells of the plurality of images;(c) identifying one or more features associated with cellular heterogeneity in one or more classified cells of the plurality of images.
6. The method of Claim 5, wherein the artificial neural network is trained according to the method of Claim 2.
7. The method of any of Claims 1 to 6, wherein identifying cellular heterogeneity involves identifying one or more different cell types, differentiation states of one or more cells, degrees of chromatin organisation, subcellular I cellular features, one or more features responsive to drug treatment, one or more features indictive of an extent of infection or disease, and I or one or more features associated with cancer cell pluripotency.
8. The method of any of Claims 1 to 7, wherein the acquired images are super resolution images or non-super resolution images.
9. The method of Claim 8, wherein the super resolution images are acquired by Singlemolecule localization microscopy (SMLM), Stochastic Optical Reconstruction Microscopy (STORM), Photoactivated Localization Microscopy (PALM), DNA points accumulation for imaging in nanoscale topography (DNA-PAINT) or a Stimulated emission depletion (STED) microscopy.
10. The method of any of Claims 1 to 9, further comprising: generating a further plurality of images of one or more cells based on one or more of the acquired images.11 . The method of Claim 10, wherein the simulated images are generated from one or more randomly selected image of the plurality of images of step (a), and wherein a generative Al approach is used to generate the simulated images.
12. The method of any of Claims 1 to 11 , wherein the feature or features are indicative of one or more cell types and / or pathologies, one or more differentiation states of a cell, chromatin organisation in a cell, transcriptional activity, subcellular / cellular characteristics of a cell, a type of cancer and / or cancer progression of a cell, precancerous characteristic of a cell (e.g. pre-leukemic characteristic), pluripotency of a cell, an extent of disease progression of a cell, response to drug treatment of a cell, and I or the viral infection level of a cell.
13. The method of any of Claims 1 to 12, wherein the feature is a biological feature, an endogenous feature of a cell, or an exogenous feature of a cell.
14. The method of any of Claims 1 to 13, wherein one or more molecule of interest in the population of cells are labelled prior to image acquisition; optionally wherein the labelling involves immunolabelling, click chemistry, DAPI staining and / or binding of one or a plurality of fluorophores to the one or more molecules of interest.
15. The method of Claim 14, wherein labelling involves binding of one or a plurality of fluorophores to the one or more molecules of interest.
16. The method of Claim 14, wherein labelling involves binding of a plurality of fluorophores wherein different ones of the fluorophores have different emission wavelengths and wherein the different ones of the fluorophores are bound to different molecules of interest.
17. The method of any of Clauses 1 to 16, wherein identifying one or more features associated with cellular heterogeneity in one or more classified cells further comprises: identifying an area of one or more images used for classification by applying a primary attribution algorithm; and determining one or more features of interest from the identified area of the one or more images.
18. The method of any preceding claim, wherein the population of one or more cells comprises, consists essentially of or consists of a single cell type.
19. The method of any preceding claim, wherein the population of one or more cells comprises multiple cell types.
20. The method of any preceding claim, wherein identifying cellular heterogeneity involves identifying one or more features indicative of presence of premalignant cells, leukemia or premalignant leukemia cells; particularly premalignant acute myeloid leukemia cells.21 . The method of any preceding claim, further comprising:generating a further plurality of images of one or more cells based on one or more of the acquired images.
22. The method of Claim 21 , wherein the further plurality of images are not super resolution images.
23. The method of Claim 21 or Claim 22, wherein the further plurality of images are obtained from diffraction limited microscopy, such as confocal microscopy images or widefield microscopy images.
24. The method of Claim 21 , wherein the further plurality of images are super resolution images.
25. The method of Claim 21 , wherein the further plurality of images are simulated images.
26. The method of Claim 25, wherein the simulated images are generated from one or more randomly selected image of the plurality of images of step (a).
27. The method of any of Claims 21 to 26, wherein the artificial neural network is trained using a training set comprising one or more of the further plurality of images.
28. The method of Claim 25, Claim 26 or Claim 27 when dependent on Claim 25 or Claim 26, wherein a generative Al approach is used to generate the simulated images.
29. The method of Claim 28, wherein the generative Al is a generative adversarial network (GAN) or a Variational autoencoder (VAE).
30. The method of Claim 29, wherein the trained artificial neural network is trained using simulated images based on randomly selected images of the plurality of images of step (a).31 . The method of any of Claims 21 to 30, wherein the number of generated images is more than about 1000, more than about 3000, more than about 5000, more than about 7000, more than about 9000, or more than about 11000.
32. The method of any of Claims 21 to 31 , wherein the number of generated images is less than about 20000, less than about 18000, less than about 16000, or less than about 14000.
33. The method of any of Claims 21 to 32, wherein the number of generated images is between about 1000 and 20000, between about 3000 and 18000, between about 5000 and 16000; or between about 7000 and 14000.
34. The method of any preceding claim, wherein the training set includes up to about 90%, up to about 80%, up to about 70%, up to about 60% or up to about 50% of the acquired plurality of images.
35. The method of any preceding claim, wherein the training set includes at least about 10%, at least about 20%, at least about 30%, at least about 40% or at least about 50% of the acquired plurality of images.
36. The method of any preceding claim, wherein the training set includes between about 10% and 90%, between about 20% and 80%, between about 30% 70%, or between about 40% and 60% of the acquired plurality of images.
37. The method of any preceding claim, wherein the test set includes up to about up to about 50%, up to about 40%, up to about 30%, up to about 20%, or up to about 10% of the acquired plurality of images; and / or at least about 5%, at least about 10%, at least about 15%, or at least about 20% of the acquired plurality of images; and / or between about 5% and 50%, between about 10% and 40% or between about 15% and 30% of the acquired plurality of images.
38. The method of any preceding claim, wherein the images of the training set are selected by random sampling of the acquired plurality of images.
39. The method of any preceding claim, wherein the training set includes at least one lower magnification image that frames one or more cells, and at least one higher magnification image that enlarges a part of one or more cells.
40. The method of any preceding claim, wherein supervised machine learning is used to identify a feature associated with cellular heterogeneity in one or more classified cells from one or more images of the training set.
41. The method of any preceding claim, wherein applying the trained artificial neural network to classify one or more cells in one or more images is based on a classification threshold, wherein the classification threshold is defined as a minimum required proportion of the images that are correctly classified using the trained artificial neural network.
42. The method of Claim 41 , wherein a classification threshold is determined by optimising true positive rate and false positive rate from a receiver operating characteristic (ROC) curve.
43. The method of Claim 41 or Claim 42, wherein the classification threshold is at least about 0.5, at least about 0.
6. at least about 0.7, at least about 0.8 or at least about 0.9.
44. The method of any of Claims 41 to 43, wherein the classification threshold is at least about 0.5, at least about 0.55, at least about 0.6, or at least about 0.65.
45. The method of any of Claims 41 to 44, wherein the classification threshold is less than about 1 , less than about 0.9, less than about 0.8, less than about 0.7, or less than about 0.6.
46. The method of any of Claims 41 to 45, wherein the classification threshold is between about 0.5 and 1 , between about 0.5 and 0.9, between about 0.5 and 0.8, between about 0.5 and 0.7, or between about 0.5 and 0.6.
47. The method of any of Claims 41 to 46, wherein the classification threshold is defined as an area under the ROC curve.
48. The method of Claim 47, wherein the area under the ROC curve is at least 0.7, at least about 0.8, or at least 0.9.
49. The method of any preceding claim, wherein identifying one or more features associated with cellular heterogeneity in one or more classified cells further comprises: identifying an area of one or more images used for classification by applying a primary attribution algorithm; and determining one or more features of interest from the identified area of the one or more images.
50. The method of Claim 49, wherein the primary attribution algorithm assigns positive and negative attribution scores to areas of the one or more images.51 . The method of Claim 49 or Claim 50, wherein the feature or features of interest are determined from areas of the one or more images with positive attribution scores.
52. The method of any of Claims 49 to 51 , wherein the primary attribution algorithm is an occlusion based algorithm.
53. The method of any of Claims 49 to 52, wherein the primary attribution algorithm is validated by a class activation mapping method.
54. A system for identifying cellular heterogeneity within a population of cells, comprising: a database library that includes a training set comprising of one or more of the acquired images, a validation set comprising one or more of the acquired images that are not within the training set and a test set comprising one or more of the acquired images that are not within the training set or the validation set; a processor in communication with the database, the processor being adapted to:receive and provide the test set to a trained artificial neural network; apply the trained artificial neural network to classify one or more cells in one or more images of the test set; and identify one or more features associated with cellular heterogeneity in one or more classified cells of the one or more images of the test set.
Citation Information
Patent Citations
Heterogeneity cell detection method and device based on deep learning
CN116309354A