Use of machine learning and / or neural networks to validate stem cells and derivatives thereof for use in cell therapy, drug creation and diagnosis
By employing machine learning and deep neural networks to analyze cell images, the method addresses the limitations of current cell evaluation techniques, providing a non-invasive, efficient, and accurate means to predict cell characteristics and develop release criteria for cell therapies.
Patent Information
- Application Number
- JP2025055757
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-03-16
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-19
AI Technical Summary
Current methods for evaluating cell health and functionality in cell therapy, drug discovery, and drug toxicity testing are limited by variability, destructiveness, cost, and inability to provide comprehensive, non-invasive assessments across the entire spectrum of drug screening and toxicity testing.
A method using machine learning and deep neural networks to non-invasively predict the characteristics of cells and their derivatives by training models with cell images and identifying features that correlate with cell function, identity, and health, allowing for the creation of release criteria for clinical cell preparations.
Enables accurate, non-invasive, and cost-effective prediction of cell characteristics, improving the efficiency of cell therapy development and drug testing while reducing the risk of contamination and cell damage.
Smart Images

Figure 2025092640000023 
Figure 2025092640000024 
Figure 2025092640000025
Abstract
Description
Technical Field
[0001] Field of the Invention The disclosed embodiments generally relate to the predictability of cell function and health, and more particularly to the use of machine learning and / or neural networks to validate stem cells and their derivatives for use in cell therapy, drug discovery, and drug toxicity testing.
Background Art
[0002] Background Many biological and clinical processes are facilitated by the provision of cells. These processes include cell therapy (e.g., stem cell therapy), drug discovery, and testing of the toxic effects of compounds. Assays for cell death, cell proliferation, cell function, and cell health are very widely performed in many fields of biological and clinical research. Several types of assays are used to evaluate cell health, death, proliferation, and functionality, including but not limited to staining with dyes, antibodies, and nucleic acid probes. For toxicity assessment and cell maturity, barrier function assays, such as trans-epithelial permeability (TEP) and trans-epithelial electrical resistance (TER), provide useful criteria. Other selected functional assays include electron microscopy (EM); gene (DNA or RNA) expression; or techniques using electrophysiological recordings, or the secretion of specific cytokines, proteins, growth factors, or enzymes; the rate and volume of fluid / small molecule transport from one side of the cell to the other; the phagocytic ability of the cell; techniques for evaluating immunohistochemical (IHC) patterns, etc., but are not limited to these. Some of these assays can be used as "release" criteria for verifying the functionality of cell therapy products prior to transplantation.
[0003] However, existing release assays used in cell therapy have significant limitations. Some of the above-described techniques, such as IHC and EM, are not quantitative. Other release assays have been found to have a high degree of variability (e.g., cytokine release, TER), or to be too destructive (e.g., gene expression and phagocytosis) or too expensive (cytokine release and gene arrays). Also, current methods require, at a minimum, opening the cell culture dish to sample the medium and, at most, completely destroying the cells being measured. In some cases, these assays cannot be performed across the entire spectrum of drug screening and toxicity testing and often do not provide comprehensive information regarding cell health and functionality. This means that it is not possible to long-term evaluate cells in their current state as they grow, develop, and differentiate, for example, prior to administration to a patient or to test drugs or toxic compounds without disturbing those cells. More specifically, many of the assays currently in use increase the potential for contamination and / or render the cells being examined unusable for any further evaluation / implantation, are not compatible with high-throughput, and do not provide a complete overall picture of cell health.
[0004] Biologists can often predict whether certain types of cells (e.g., retinal pigment epithelial (RPE) cells) are functional just by looking at them under brightfield / phase contrast microscopy. This approach is non-invasive, versatile, and relatively inexpensive. Despite these advantages, human-based image analysis has drawbacks, including but not limited to, specimen extraction bias, difficulty in directly correlating visual data with function, and difficulty in inferring causality.
[0005] Therefore, there is a need in the art for automation of analysis that can address the above limitations of current approaches in order to deliver cell and tissue therapies to patients and to more efficiently discover new drugs and potentially toxic compounds.
Summary of the Invention
Means for Solving the Problems
[0006] Gist of the Invention The objectives and advantages of the embodiments illustrated below will be shown and will become apparent from the following description. Further advantages of the illustrated embodiments can be realized and achieved by the devices, systems and methods specifically pointed out in the specification and claims, as well as from the accompanying drawings.
[0007] According to an aspect of the present disclosure, there is provided a method for non-invasively predicting the characteristics of one or more cells and cell derivatives. The method includes training a machine learning model using at least one of a plurality of training cell images representing a plurality of cells and data identifying the characteristics of the plurality of cells. The method further includes receiving at least one test cell image representing at least one test cell to be evaluated, wherein the at least one test cell image is non-invasively acquired and is based on absorbance as an absolute scale of light, and providing the at least one test cell image to the trained machine learning model. Using machine learning based on the trained machine learning model, the characteristics of the at least one test cell are predicted. The method further includes creating release criteria for a clinical cell preparation by the trained machine learning model based on the predicted characteristics of the at least one test cell.
[0008] In an embodiment, the machine learning can be performed using a deep neural network, and the method can further include segmenting the image of the test image into individual cells by the deep neural network.
[0009] Furthermore, in an embodiment, the machine learning can be performed using a deep neural network, and the method can further include classifying the test cells based on the characteristics.
[0010] In a further embodiment, the prediction step may be based on a classification step and may further include determining at least one of cell identity, cell function, effect of the delivered drug, disease state, and technical replicates or similarity to previously used samples.
[0011] In an embodiment, the prediction of the characteristics of the test cells can be performed on either a single cell or a plurality of cells in one field of view in at least one test cell image.
[0012] In an embodiment, the method may further include visually extracting at least one feature from the test cell image. The step of training the machine learning model can be performed using the at least one extracted feature, and the prediction step includes identifying at least one feature of the test cells using the trained machine learning model trained using the at least one extracted feature, and predicting the characteristics of the test cells using the at least one identified feature.
[0013] Furthermore, in an embodiment, the test cell image is acquired using quantitative bright-field absorbance microscopy (QBAM).
[0014] Furthermore, in an embodiment, the method may further include receiving at least one microscopic image captured by a microscope and converting the pixel intensity of the at least one microscopic image into an absorbance value.
[0015] In an embodiment, the method may further include at least one of calculating the absorbance reliability of the absorbance value, establishing the balance of the microscope by benchmarking, and filtering the color when acquiring the image.
[0016] In addition, in an embodiment, the first image processing of the test image can be performed by a deep neural network.
[0017] In an embodiment, machine learning can be performed using a deep neural network, and the method may further include a step of segmenting an image of at least one test image into individual cells by the deep neural network, and the features are visually extracted from the segmented individual cells.
[0018] In an embodiment, the prediction of the characteristics of the at least one test cell may include at least one of a prediction of the polarized secretion of transepithelial resistance (TER) and / or vascular endothelial growth factor (VEGF), a prediction of the function of the at least one test cell, a prediction of the maturity of the at least one test cell, a prediction of whether the at least one test cell is derived from a donor that has been identified, and a determination of whether the at least one test cell is an outlier with respect to a known classification.
[0019] In a particular embodiment, the creation of release criteria for a clinical cell preparation may further include the creation of release criteria for drug discovery or drug toxicity by a trained machine learning model.
[0020] In addition, in an embodiment, the plurality of cells and test cells may include at least one of embryonic stem cells (ESCs), induced pluripotent stem cells (iPSCs), neural stem cells (NSCs), retinal pigment epithelial stem cells (RPESCs), mesenchymal stem cells (MSCs), hematopoietic stem cells (HSCs), and cancer stem cells (CSCs).
[0021] In an embodiment, the first plurality of cells and test cells may be derived from at least one of a plurality of ESCs, iPSCs, NSCs, RPESCs, MSCs, or HSCs or any cells derived therefrom.
[0022] Furthermore, in an embodiment, the first plurality of cells and test cells may include the major cell types derived from human or animal tissues.
[0023] In addition, in embodiments, the identified and predicted characteristics may include at least one of physiological characteristics, molecular characteristics, cellular characteristics, and / or biochemical characteristics.
[0024] In embodiments, at least one of the extracted features may include at least one of a cell boundary, a cell shape, and a plurality of texture metrics.
[0025] Furthermore, in embodiments, the plurality of texture metrics may include features within a plurality of cells.
[0026] Further, in embodiments, the plurality of cell images and test cell images may include fluorescence images, chemiluminescence images, radiographic images, or bright-field images.
[0027] In embodiments, the method further may include determining whether a test image of at least one test image is a large image, dividing the large image into at least two tiles, individually providing each of the tiles to a trained machine model, combining the outputs of the processing by the trained machine model associated with each of the tiles, and providing the combined output to predict the characteristics of the test cell to obtain a single output corresponding to the large image.
[0028] In a further aspect of the present disclosure, a computing system for performing the disclosed method is provided.
[0029] These and other features of the systems and methods of the present disclosure will become readily apparent to those skilled in the art from the following detailed description of the preferred embodiments taken in conjunction with the drawings.
[0030] This disclosure will become more fully understood from the detailed description and the accompanying drawings. These accompanying drawings illustrate one or more embodiments of the disclosure and, together with the specification, serve to explain the principles of the disclosure. Throughout the drawings, the same reference numbers are used to refer to the same or similar elements of an embodiment, wherever possible.
Brief Description of the Drawings
[0031]
Figure 1
[0032]
Figure 2
[0033]
Figure 3-1
Figure 3-2
[0034]
Figure 4
[0035]
Figure 5
[0036]
Figure 6
[0037]
Figure 7
[0038]
Figure 8
[0039]
Figure 9
[0040]
Figure 10
[0041]
Figure 11
[0042]
Figure 12
[0043]
Figure 13
[0044]
Figure 14
[0045] The following exemplified embodiments are merely illustrative and, as will be recognized by those skilled in the art, can be embodied in various forms, so those exemplified embodiments are in no way limited to what is illustrated. Thus, any structural and functional details disclosed herein should not be construed as limiting, but rather should be understood merely as a representation for teaching those skilled in the art to use the claims and the contemplated embodiments in various ways. Further, the terms and phrases used herein are not intended to be limiting but rather are intended to provide an understandable description of the exemplified embodiments. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the exemplified embodiments, but the exemplary methods and materials will be described hereinafter.
[0046] It should be noted that, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, a reference to "a stimulus" includes a plurality of such stimuli, and a reference to "the signal" includes references to one or more signals and their equivalents known to those skilled in the art.
[0047]
[0048] The exemplary embodiments discussed below should be recognized as software algorithms, programs or code residing on a computer-usable medium having control logic for enabling execution in a machine having a computer processor. The machine typically comprises a storage device configured to provide an output from the execution of a computer algorithm or program.
[0049] As used herein, the term "software" is meant to be synonymous with any code or program that can exist on the processor of a host computer, regardless of whether its execution takes place in hardware, in firmware, or as a downloadable software computer product on a disk, on a storage device, or from a remote machine. The embodiments described herein include such software for executing the equations, relationships, and algorithms described below. Those skilled in the art will recognize additional features and advantages of the exemplary embodiments based on what is described above. Accordingly, the exemplary embodiments should not be particularly shown and particularly described as being limited to those shown by the appended claims.
[0050] In an exemplary embodiment, computer system components may comprise "modules" configured to perform and operative to perform certain operations described below herein. Accordingly, the term "module" should be understood to encompass a tangible entity that is physically constructed (e.g., physically incorporated), permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a particular manner and to perform certain operations described herein.
[0051] FIG. 1 illustrates an overall overview of one operating environment 100 according to one embodiment of the present disclosure. In particular, FIG. 1 illustrates an information processing system 102 that can be used in an embodiment of the present disclosure. The processing system 102, disclosed in more detail below, can be used to verify the function of grafts in clinical biomanufacturing (including non-invasively predicting tissue function and donor identity of cells). The processing system 102 is compatible with advances in developmental biology and regenerative medicine that help create cell-based therapies for treating retinal degeneration, neurodegeneration, heart disease, and other diseases by replacing damaged or degenerated native tissue with new functional implants developed in vitro. Induced pluripotent stem cells (iPSCs) have expanded the potential of cell therapies that enable autologous tissue transplantation. The methods of the present disclosure supported by the processing system 102 are performed non-invasively, eliminating the requirements for trained users, enabling high-throughput, reducing costs, and shortening the time required to make appropriate predictions (e.g., regarding tissue function and donor identity of cells), thereby overcoming barriers to large-scale biomanufacturing.
[0052] The information processing system 102 shown in FIG. 1 is merely an example of a suitable system and is not intended to limit the scope of use or functionality of the embodiments of the present disclosure described below. The information processing system 102 of FIG. 1 can perform and / or carry out any of the functionality shown below. Any appropriately configured processing system can be used as the information processing system 102 in the embodiments of the present disclosure.
[0053] As illustrated in FIG. 1, components of the information processing system 102 can include, but are not limited to, one or more processors or processing devices 104, a system memory 106, and a bus 108 that couples various system components including the system memory 106 to the processor 104.
[0054] Bus 108 corresponds to one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor bus or local bus, using any of various bus architectures. By way of example, and not limitation, such architectures include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Extended ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0055] System memory 106, in one embodiment, includes a machine learning module 109 configured to perform one or more embodiments discussed below. For example, in one embodiment, machine learning module 109 is configured to generate an output vector based on an input array of measurements using a machine learning prediction model by applying selected machine learning. In another embodiment, machine learning module 109 is configured to generate an output based on an input array of images using a deep neural network model (these are discussed in more detail below). FIG. 1 shows machine learning module 109 resident in main memory, but note that machine learning module 109 can be resident within processor 104, can be a separate hardware component corresponding to multiple information processing systems and / or processors, and / or can be distributed across multiple information processing systems and / or processors.
[0056] System memory 106 may also include computer system readable media in the form of volatile memory such as random access memory (RAM) 110 and / or cache memory 112. Information processing system 102 may further include other removable / non-removable volatile / non-volatile computer system storage media. By way of example only, storage system 114 may be provided for reading from and writing to non-removable or removable non-volatile media, such as one or more solid state disks and / or magnetic media (commonly referred to as a “hard drive”). A magnetic disk drive for reading from and writing to a removable non-volatile magnetic disk (e.g., a “floppy (registered trademark) disk”) as well as an optical disk drive for reading from or writing to removable non-volatile optical disks (e.g., CD-ROM, DVD-ROM or other optical media) may be provided. In such cases, each may be connected to bus 108 by one or more data media interfaces. Memory 106 may include at least one program product having a set of program modules configured to perform the functions of embodiments of the present disclosure.
[0057] Program / utility 116 having a set of program modules 118 may be stored in memory 106 by way of example and not limitation, as well as an operating system, one or more application programs, other program modules, and program data. Each of the operating system, one or more application programs, other program modules, and program data, or some combination thereof, may include an implementation of a network usage environment. Program modules 118 generally perform the functions and / or methods of embodiments of the present disclosure.
[0058] The information processing system 102 can also communicate with one or more external devices 120 (e.g., keyboard, pointing device, display 122, etc.); one or more devices that enable a user to interact with the information processing system 102; and / or any device that enables the information processing system 102 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication can be performed via the I / O interface 124. Nevertheless, the information processing system 102 can communicate with one or more networks (e.g., local area network (LAN), general wide area network (WAN), and / or public network (e.g., Internet)) via the network adapter 126. As depicted, the network adapter 126 communicates with other components of the information processing system 102 via the bus 108. Other hardware components and / or software components can also be used with the information processing system 102. Examples include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archive storage systems.
[0059] Various embodiments of the present disclosure relate to a computational framework that uses novel neural networks and / or machine learning algorithms to analyze images of stem cells and / or to analyze multiple cell types derived from these stem cells. At least in some embodiments, by using neural networks and / or machine learning, one or more physiological and / or biochemical functions of cells can be automatically predicted.
[0060] The images input to the information processing system 102 can be acquired using an external device 120 such as a bright-field microscope. The information processing system 102 can automatically process these images in real time or at a selected time point using quantitative bright-field absorbance microscopy (QBAM). QBAM uses absorbance imaging based on absorbance, which is an absolute scale of light. The QBAM method used can be implemented on any standard bright-field microscope and provides a statistically robust method for outputting reliable and reproducible image quality.
[0061] Specifically, the application of QBAM involves converting the pixels of the input image from relative intensity units to absorbance units, which are the absolute scale of light attenuation. To improve the reproducibility of imaging, QBAM calculates the statistical quantities of the image in real time as the image is captured, ensuring that the absorbance values measured at each pixel have a 95% confidence level of 10 milli-absorbance units (mAU).
[0062] The unprocessed pixel intensities can vary depending on the microscope setup and settings (e.g., non-uniform illumination, bulb intensity and spectrum, camera, etc.), even when the image is captured on the same microscope, making it difficult to compare images. Therefore, QBAM offers advantages over methods that use unprocessed pixel intensities. Thus, the method of the present disclosure can provide automation, conversion from pixel intensity to absorbance value, calculation of absorbance confidence, and establishment of microscope balance through the use of benchmarking to maximize the quality of the image data captured using QBAM.
[0063] In scenarios where images are processed using transmittance values, such as when using histological examinations, the calculated statistical quantities can be modified to generate reproducible tissue section images. Since the calculation of QBAM is not wavelength-specific, it can be generalized to any multispectral modality. The method of the present disclosure may be suitable for hyperspectral autofluorescence imaging that can identify cell boundaries and intracellular structures of non-pigmented cells.
[0064] The information processing system 102 may use one or more processing systems and / or software applications (including for applying machine learning and / or deep neural network processing to the image, training a model, extracting features, performing a clustering function, and / or performing a classification function) to process the received image. In an embodiment, an image output by a microscope is converted from one format (e.g., but not limited to CZI) to another format (e.g., but not limited to JPEG or TIFF) before further processing is performed. In an embodiment, a script of an application such as (but not limited to) MATLAB® that receives an image from a microscope is adapted to manage the format output by the microscope. In an embodiment, the received image and the data obtained from the received image may be used by one or more different programs (e.g., MATLAB®, Fiji by ImageJ TM , C++, and / or Python®). In an embodiment, the data may be manually transferred from one program to another program. In an embodiment, the script is adapted to enable the programs to directly transfer data to each other.
[0065] Accordingly, the disclosed method may be executed in any clinical-grade biomanufacturing environment or high-throughput screening system, and the disclosed method requires only a brightfield microscope, a digital camera, and a desktop computer for analysis.
[0066] This description presents a set of data-driven modeling tools that can be used to evaluate the functions of human stem cells (e.g., iPSCs and retinal pigment epithelium (RPE) cells). At least in some embodiments, this image analysis is used to create lot and batch release criteria for clinical preparations of iPSCs and RPEs. In certain embodiments, the disclosed system can be utilized to deliver cell and various types of tissue therapies to patients in a more efficient manner. Further, the disclosed approach can be used to 1) distinguish differences between different subtypes of cells, thus providing the possibility of creating or optimizing the production of specific cell and tissue types using stem cells; 2) distinguish differences between healthy cells and diseased cells, thereby providing the possibility of discovering the underlying pathogenesis of drugs or diseases and improving the health status of diseased cells; 3) distinguish differences between cells treated with drugs and toxins, thereby providing the possibility of determining the efficacy of drugs, drug side effects, and the effects of toxins and their combinations on cells. Each of these specific methods can be performed without disturbing the cells being examined at a particular point in time and / or over a period in a longitudinal study.
[0067] Figures 2A and 2B illustrate selected machine learning frameworks and deep neural network frameworks that can be used individually or together to predict the function, identity, disease state, and health of cells and their derivatives, according to certain embodiments of the present disclosure.
[0068] In the embodiment illustrated in Figure 2A, the machine learning prediction model 204 is trained to correlate input data 202 by a data-driven statistical analysis approach to generate a new interpretable model. In certain embodiments, the input data 202 can include an input array of measurements representative of at least one physiological parameter, molecular parameter, cellular parameter, and / or biochemical parameter of a plurality of stem cells or cell types derived from a plurality of stem cells.
[0069] The input data 202 can be obtained from the extracted visual features of individual cells extracted from the QBAM image, for example, using a web image processing pipeline (WIPP). By using the extracted visual features, a selected machine learning algorithm can be trained to predict various tissue characteristics (including function, the identity of the donor from which the cells are derived, and developmental outliers (the appearance of abnormal cells)). Then, by using the selected machine learning algorithm, decisive cell features that can contribute to the prediction of tissue characteristics can be identified.
[0070] In various embodiments, the plurality of stem cells can include at least one of embryonic stem cells (ESC), induced pluripotent stem cells (iPSC), neural stem cells (NSC), retinal pigment epithelial stem cells (RPESC), retinal pigment epithelial cells derived from induced pluripotent stem cells (iRPE), mesenchymal stem cells (MSC), hematopoietic stem cells (HSC), cancer stem cells (CSC), or any cell type derived therefrom, of human or animal origin. In some embodiments, the input data 202 can include an input array of measurements representing at least one physiological parameter, molecular parameter, cellular parameter, and / or biochemical parameter of a plurality of major cell types derived from human tissue or any animal tissue.
[0071] In various embodiments, the machine learning prediction model 204 can include one of a plurality of types of prediction models (dimensionality reduction / rank reduction, hierarchical clustering / grouping, vector-based regression, decision tree, logistic regression, Bayesian network, random forest, etc.). The machine learning prediction model 204 is trained on similar cell metrics and is used to provide a more robust grouping of both different patients and clones of iPSC-derived RPE tissue engineering constructs. Thus, the machine learning prediction model 204 generates an output vector 206 that represents, for example, the physiological and / or biochemical functions of the cells by selecting and ranking the individual visual parameters of the cells.
[0072] According to another embodiment of the present disclosure illustrated in FIG. 2B, a neural network framework using, for example, a deep neural network model 212 is trained using a training input array 208 that includes labeled and / or defined images and / or data. The images of the training input array 208 show the microstructure of various stem cells or stem cell-derived cell types. In one embodiment, the image of each microstructure may preferably include the microstructure of cells captured at a magnification of 2× to 200×. Additionally or alternatively, the training input array 208 may include an image array in which the ultrastructure or microstructure data of the cells that the researcher is interested in is highlighted with identifiers, and may also include data identifying the origin of the cells, the location of the cells, the cell line, the physiological data, and the biochemical characteristics for the corresponding plurality of stem cells or stem cell-derived cell types.
[0073] As described in detail below, the deep neural network model 212 can consistently and autonomously analyze images, identify features within the images, perform high-throughput segmentation of a given image, and correlate the image with identity, safety, physiological results, biochemical results, or molecular results.
[0074] Furthermore, the segmentation generated by the deep neural network model 212 can be overlaid again on the original microstructure images of the training input array 208, and then image feature extraction can be performed on those microstructure images. Various quantitative cell information (including but not limited to cell morphometry (area, perimeter, elongation, etc.), intensity (average luminance, mode, median, etc.), or texture (intensity entropy, standard deviation, uniformity, etc.)) can be calculated for individual features within the cell, for the cell itself, or in the tissue region where the stem cells or stem cell-derived cell types are distributed.
[0075] Among the advantages of the deep neural network model 212 contemplated by various embodiments of the present disclosure are the prediction reliability provided by the training input array 208 and the ability to consistently determine the complex relationships between cell images and cell identity, safety, physiological processes, and / or biochemical processes.
[0076] Then, by using the segmentation generated by the deep neural network model 212 and applying a selected machine learning method to the features extracted from these images, it is possible to identify which features in the image are correlated with cell identity, safety, physiological function, and / or biochemical processes. Further, the deep neural network model 212 can reduce the uncertainty of the training input array 208.
[0077] The key to the process provided by the deep neural network model 212 for predicting the function, identity, disease state, and health of cells and their derivatives is the ability to automatically adjust the grouping of parameters in a quantitative form based on the characterized images of the training input array 208. Various embodiments of the present disclosure relate to novel methods of cell function recognition and characterization focused on deep learning, where the recognition of various functions can be performed using only visual data. The methods disclosed below should be regarded as a non-limiting example of a method using a deep convolutional neural network for the characterization of cell functionality. Methods and systems for training the deep neural network model 212 to handle limited labeled data are disclosed herein.
[0078] When a large amount of annotated data is available, other possible frameworks used for characterizing the functionality of cells typically use pixel-level segmentation in a supervised setting. Examples of such methods / frameworks include, but are not limited to, fully convolutional networks (FCNs), U-Net, pix2pix, and their extensions. However, the FCN framework used for other tasks such as semantic image segmentation and super-resolution is not as deep as the disclosed method (depth is related to the number of layers). For example, an FCN network can receive an image as input and generate the entire image as output through four hidden layers of convolutional filters. The weights are learned by minimizing the difference between the output and the clean image.
[0079] One of the difficulties in applying deep learning methods to cell image analysis is the lack of a large amount of annotated data for training neural network models. However, humans can use their visual systems to recognize patterns from both daily natural images and microscopic images. Therefore, at a very high level, the goal is to train a model using one type of image and teach the model to automatically learn data patterns that can be used for the task of recognizing cell functionality.
[0080] In one embodiment, the deep neural network framework uses a two-step approach. In one example, images of the training input array 208 include live fluorescence microscopy images, multispectral brightfield absorption images, chemiluminescence images, radioactive images, or hyperspectral fluorescence images showing RPE with anti-ZO-1 antibody staining (white). The tight junction protein ZO-1 represents the boundaries of RPE cells. In other examples. Advantageously, based on an understanding of the cell boundaries and visual parameters (i.e., shape, intensity, and texture metrics) in the microscopic images, the deep neural network model 212 can detect the correlation of such cell boundaries and visual parameters in new input array 210 images (e.g., live fluorescence microscopy images, multispectral absorption brightfield images, chemiluminescence images, radioactive images, or hyperspectral fluorescence images of similar cells or cell-derived products). It should be noted that texture metrics can include features within multiple cells. Additionally, it should be noted that the multispectral brightfield absorption images can include phase contrast, differential interference contrast, or any other images with brightfield modalities.
[0081] Embodiments of the present disclosure use a concept known as transfer learning in the machine learning community. In other words, the disclosed method takes advantage of effectively and automatically retaining knowledge and effectively and automatically transferring knowledge from one task (microscopic images) to another related task (analysis of live absorbance images).
[0082] As described above, in various embodiments, the machine learning prediction model 204 can include multiple types of prediction models. In one embodiment, the machine learning prediction model 204 can use dimensionality reduction and / or rank reduction. FIGS. 3A - 3C illustrate the conversion from multi-dimensional data to low-dimensional data using a principal component analysis machine learning model according to an embodiment of the present disclosure.
[0083] When using factor analysis as a variable reduction method, the correlation of two or more variables can be summarized by combining the two variables into a single factor. For example, two variables can be plotted on a scatter plot. A regression line representing the general outline of the linear relationship between those two variables can be fitted (e.g., by the machine learning prediction model 204 in FIG. 2A). For example, if two variables exist, a two-dimensional plot can be made where those two variables define a line. In the case of three variables, a three-dimensional scatter plot can be determined and a plane can be fitted through the data. In the case of more than three variables, it becomes difficult to illustrate points on a scatter plot, but the analysis is performed by the machine learning prediction model 204 and a summary of the regression of the relationship of those three or more variables can be determined. In a plot that captures the principal components of two or more items, variables approximating the regression line can be defined. Data measurements from stem cell data or stem cell-derived cell type data for the new factor (i.e., represented by the regression line) can be used in future data analysis to represent the essence of those two or more items. Thus, two or more variables may be reduced to one factor, where the factor is a linear combination of those two or more variables.
[0084] The extraction of the principal components (e.g., the first component 302 and the second component 304 in FIG. 3A) can be found by determining the variance-maximizing rotation of the original variable space. For example, in the scatter plot 310 shown in FIG. 3B, the regression line 312 can be the original X-axis in FIG. 3A rotated to approximate the regression line. This type of rotation is called variance maximization. This is because the criterion (i.e., the goal) of the rotation is to maximize the variance (i.e., the spread) of the "new" variables (factors) while minimizing the variance around the new variables. It is difficult to create a scatter plot using three or more variables, but the logic of rotating the axes to maximize the variance of the new factors remains the same. In other words, the machine learning prediction model 204 continues to plot the next optimal line 312 based on the multi-dimensional data 300 shown in FIG. 3A. FIG. 3C shows the final optimal line 322, where all the variance of the data is explained by the machine learning prediction model 204.
[0085] According to another embodiment, the machine learning prediction model 204 may use a clustering approach. FIG. 4 shows an example of a hierarchical clustering method for identifying similar groups according to an embodiment of the present disclosure. The clustering approach includes cluster analysis. In the context of the present disclosure, the term "cluster analysis" encompasses several different standard algorithms and methods for grouping similar types of objects into respective categories and thus integrating the observed data into a meaningful structure. In this context, cluster analysis is a general data analysis process for sorting different objects into groups such that the degree of association between the objects is maximized when the two objects belong to the same group and minimized when they do not. Cluster analysis can be used to find the structure within the data without necessarily providing an explanation for the grouping. In other words, cluster analysis can be used to find the logical grouping within the data.
[0086] In an alternative embodiment, the clustering approach may include hierarchical clustering. The term "hierarchical clustering" involves grouping the measured results into successively less similar groups (clusters). In other words, hierarchical clustering uses some measure of similarity / distance.
[0087] Figure 4 illustrates an example of a hierarchical cluster 402. Clustering can be used to determine the safety of a cell therapy product by identifying which cell lines have mutations in cancer genes. In this case, hierarchical bootstrap clustering can be performed on a dataset of cell features obtained from the segmentation of RPE images by a convolutional neural network. Cell lines with cancerous mutations are statistically different from all other non-mutated lines (see, for example, Figure 13). In various embodiments, the clustering approach discussed above can be used to determine how similar groups, such as treatments, genes, etc., are related to each other.
[0088] According to yet another embodiment, the machine learning prediction model 204 may use vector-based regression. Refer to Figure 5, which shows an example of a vector-based regression model that the machine learning prediction model 204 according to an embodiment of the present disclosure may use. More specifically, Figure 5 shows a two-dimensional data region 500. It will be recognized that this data region may be n-dimensional in other embodiments, and that the hyperplane of an n-dimensional space is an (n - 1)-dimensional plane (for example, in three-dimensional space, the hyperplane is a two-dimensional plane). In the case of a two-dimensional data region, the hyperplane is a one-dimensional straight line in that data region.
[0089] The data region 500 has measurement value sample examples 501, 503. The training samples include samples of one class (here called "first clone" samples 501) and samples of another class (here called "second clone" samples 503). Three candidate hyperplanes labeled 502, 504, and 506 are shown. Note that the first hyperplane 502 does not separate the two classes, the second hyperplane 504 also does not separate the two classes, but has a relatively smaller margin than the third hyperplane 506, and the third hyperplane 506 separates the two classes with the largest margin. Thus, the machine learning prediction model 204 outputs 506 as the classification hyperplane, that is, the hyperplane used to classify new samples as either the first clone or the second clone.
[0090] The measured values identified as outliers can be utilized in regression analysis to analyze specific parameters, relationships between parameters, etc. The dimension is defined as a set of samples for which the dot product of the vectors is always constant. The vectors are usually selected such that the distance between the samples is maximized. In other words, a penalty is imposed on the candidate vectors for proximity to the collected data set. In various embodiments, this penalization, as well as the distance at which such penalization occurs, can be configurable parameters of the machine learning prediction model 204.
[0091] The lower part of FIG. 5 illustrates the hyperplane 506 compared to other parallel hyperplanes 512 and 514. The hyperplanes 512 and 514 separate the data of the two classes such that the distance between them is as large as possible. The third hyperplane 506 is located in the middle of the hyperplanes 512 and 514.
[0092] The prediction models shown in FIGS. 3-5 are merely examples of suitable machine learning methods that the machine learning prediction model 204 may use, and are not intended to limit the scope of use or functionality of the embodiments of the present disclosure described above. Other machine learning methods that the machine learning prediction model 204 may use include, but are not limited to, partial least squares regression, local partial least squares regression, orthographic projection onto latent structures, three-pass regression filters, decision trees with recursive feature elimination, Bayesian linear and logistic models, Bayesian ridge regression, and the like.
[0093] FIG. 6 illustrates an exemplary fully-connected deep neural network (DNN) 600 that may be executed by the deep neural network model 212 according to an embodiment of the present disclosure. The DNN 600 includes a plurality of nodes 602 organized into an input layer 604, a plurality of hidden layers 606, and an output layer 608. Each of the layers 604, 606, 608 is connected by a node output 610. The number of nodes 602 shown in each layer is meant to be exemplary and is not to be understood as limiting in any way. Additionally, although the illustrated DNN 600 is shown as fully-connected, the DNN 600 may have other arrangements.
[0094] As an overview of the DNN 600, an image 603 to be analyzed may be input to the nodes 602 of the input layer 604. Each of the nodes 602 may correspond to a mathematical function with adjustable parameters. All of the nodes 602 may be the same scalar function that differs only, for example perhaps, by different parameter values. Alternatively, the various nodes 602 may be different scalar functions depending on the layer position, input parameters, or other distinguishing features. By way of example, the mathematical function may take the form of a sigmoid function. It is understood that other function forms may also be used additionally or optionally. Each of those mathematical functions may be configured to receive one input or a plurality of inputs, and to calculate or compute a scalar output from that one input or those plurality of inputs. By way of example of a sigmoid function, each node 602 may calculate the non-linearity of the sigmoid curve of the weighted sum of its inputs.
[0095] Thus, the nodes 602 in the input layer 604 receive the microscopic images 603 of the cells and then generate node outputs 610, which are continuously delivered through the hidden layer 606. The node outputs 610 of the input layer 604 are directed towards the nodes 602 of the first hidden layer 606, the node outputs 610 of the first hidden layer 606 are directed towards the nodes 602 of the second hidden layer 606, and so on. Finally, the nodes 602 of the last hidden layer 606 can be delivered to the output layer 608, which can then output a prediction 611, for example, regarding specific cell characteristics (e.g., TER) or cell identity.
[0096] Prior to the runtime use of the DNN 600, the DNN 600 can be trained using labeled or transcribed data. For example, during training, the DNN 600 is trained using a set of labeled / defined images of the training input array 208. As shown, the DNN 600 is considered to be "fully connected". This is because the node outputs 610 of each node 602 in the input layer 604 and the hidden layer 606 are connected to the inputs of all nodes 602 in either the next hidden layer 606 or the output layer 608. Thus, each node 602 receives input values from the immediately preceding layers 604, 606, except for the nodes 602 in the input layer 604 that receive the input data 603. Embodiments of the present disclosure are not limited to feed-forward networks, and the use of recurrent neural networks is equally contemplated, where at least some layers feed data back to earlier layers.
[0097] In at least one embodiment, the deep neural network model 212 may use a convolutional neural network. FIG. 7 is a schematic diagram of an exemplary convolutional neural network model architecture according to an embodiment of the present disclosure. The architecture in FIG. 7 shows a plurality of feature maps, also known as activation maps. In one illustrative example, if the target image 702 to be classified (e.g., one of the images of the new input array 210) is a JPEG image having a size of 224×224, a representative array of that image may be 224×224×3 (where "3" refers to RGB values). The corresponding feature maps 704-718 may be represented by the following arrays 224×224×64, 112×112×128, 56×56×256, 28×28×512, 14×14×512, 7×7×512, 1×1×4096, 1×1×1000, respectively.
[0098] Furthermore, as described above, a convolutional neural network includes a plurality of layers, one of which is a convolutional layer that performs convolution. Each convolutional layer acts as a filter for filtering the input data. At a high level, the CNN takes the image 702 and passes it through a series of convolutional layers, non-linear layers, pooling (downsampling) layers, and fully connected layers to obtain an output. The output can be a single class or a class with a certain probability that best explains the image.
[0099] Convolution involves the generation of an inner product based on a filter and input data. Conventionally, immediately after each convolutional layer, a non-linear layer (or activation layer), such as a ReLU (Rectified Linear Unit) layer, is applied. The purpose of the non-linear layer is basically to introduce non-linearity into a system that has just computed linear operations (e.g., element-wise multiplication and addition) during the operation by the convolutional layer. After several ReLU layers, the CNN may have one or more pooling layers. The pooling layer is also referred to as a downsampling layer. There are also several layer options in this category, such as max pooling. This max pooling layer basically takes a filter (usually of size 2×2) and a stride of the same length, and the max pooling layer is applied to the input volume and outputs the maximum number in each sub-region where the filter is convolved.
[0100] Each fully connected layer receives an input volume (e.g., its output is the output of the immediately preceding convolutional layer, ReLU layer, or pooling layer) and outputs an N-dimensional vector (where N is the number of classes that the learning model has to choose from). Each number in this N-dimensional vector represents the probability of a certain class. The fully connected layer processes the output of the previous layer (corresponding to the activation map of high-level features) and determines which features are most correlated with a specific class. For example, a specific output feature from the previous convolutional layer may indicate whether a specific feature in that image represents an RPE cell, and by using such features, the target image can be classified as "RPE cell" or "non-RPE cell".
[0101] Furthermore, an exemplary CNN architecture can clearly model a bipartite graph representation (BGL) by having a softmax layer together with the last fully connected layer, and by using it, the CNN can be optimized using global back-propagation.
[0102] More specifically, the exemplary architecture of the CNN network shown in FIG. 7 includes a plurality of convolutional and ReLU layers 720, max pooling layer 722, fully connected layer and ReLU layer 724, and softmax layer 726. In one embodiment, the CNN network 700 may include 139 layers and 42 million parameters.
[0103] FIG. 8 illustrates a stem cell image that can be used by a conventional machine learning framework and / or a deep neural network framework to measure the function of stem cells, according to an embodiment of the present disclosure. In this embodiment, iPSC-RPE cells of two different cell lines are seeded in each well of a 12-well dish. It should be noted that this RPE cell is essential for photoreceptor development and function and requires functional primary cilia for complete maturation. One set of cells is left untreated, and the other two sets are manipulated at the maturation stage of iPSC-RPE differentiation. Another set of cells is treated with aphidocolin, a tetracyclic antibiotic that increases ciliogenesis and promotes RPE differentiation by blocking the transition of cells from G1 to S after 10 days of culture.
[0104] Treat another set of seeded cells with HPI-4, an AAA+ ATPase dynein motor inhibitor that functions by blocking ciliary protein transport and inhibiting ciliary function, thus inhibiting RPE differentiation and providing a preferred negative control. According to certain embodiments of the present disclosure, measurements are taken at six different time points (e.g., weeks 2-7) during the maturation stage of iPSC-RPE differentiation. Such measurements include, but are not limited to, TER, cytokine secretion profiles, and phagocytic capacity. It should be noted that one of the most important functions of RPE cells is the phagocytosis of photoreceptors that shed their outer segments. FIG. 8 illustrates RPE cells 802 treated with aphidicolin, untreated RPE cells 804 (control group), and RPE cells 806 treated with HPI-4. A plurality of images 802-806 are used as images of the new input array 210 to be analyzed. The CNN network 700 is configured to predict TER measurement values for the plurality of images. In one embodiment, the plurality of input images can include approximately 15,000 images.
[0105] FIG. 9 illustrates the predicted results 900 of TER provided by the CNN model according to certain embodiments of the present disclosure. In FIG. 9, the horizontal axis 906 represents the actual TER measurement values, and the vertical axis 908 represents the TER values predicted by the CNN network 700. Results for all three sets of grown RPE cells and for all measurements taken at different time points are shown. In this embodiment, the release criterion is configured as a TER of 400 Ohm*Cm2. In FIG. 9, region 902 includes all cells that were correctly accepted, and region 904 represents all cells that were correctly rejected by the CNN model 700. In other words, in this case, the CNN model 700 is a perfect predictor that can predict the average cell resistance (TER) with 100% sensitivity and 100% specificity.
[0106] Figure 10 illustrates that there is no direct correlation between the bulk absorbance measurement and the TER. More specifically, Figure 10 shows a plot 1002 of absorbance vs. TER. In Figure 10, the horizontal axis 1004 represents the actual TER measurement values, and the vertical axis 1006 represents the absorbance values. Plot 1002 clearly illustrates that there is no direct correlation between the bulk absorbance measurement and the TER measurement values. Therefore, the multiple visual parameters (e.g., spatial data, texture data, and geometric data) used by the CNN model 700 are necessary for the prediction ability.
[0107] Figure 11 illustrates the segmentation based on the similarity approach used by the CNN model according to an embodiment of the present disclosure. In Figure 11, the first image 1102 represents an image of the training input array 208 including a confocal fluorescence microscope image showing RPE having anti-ZO-1 antibody staining (white). The tight junction protein ZO-1 represents the boundary of RPE cells. The second image 1104 represents the first image 1102 segmented by the CNN model 700. Further, in this embodiment, the image of the new input array 210 to be analyzed includes the live multispectral absorption image represented in Figure 11 by the third image 1106. The fourth image 1108 represents the third image 1106 segmented by the deep neural network model 212. It should be noted that the second image 1104 can be verified by manually correcting the clearly visible cell boundaries in the first image 1102, but for the fourth image 1108, although manual correction is not impossible, it is quite difficult. Advantageously, based on the understanding of cell boundaries and visual parameters (e.g., shape, intensity, and texture metrics), the deep neural network model 212 can detect cell boundaries and correlate visual parameters within similar multispectral absorption images of the new input array 210.
[0108] FIG. 12 illustrates a comparison between principal component analysis performed using only molecular data and physiological data, and image analysis performed using only visual data, according to an embodiment of the present disclosure. Each combination of shape-shading in FIG. 12 shows different clones made from one of three donors (herein named "Donor 2", "Donor 3", and "Donor 4"). By performing PCA dimensionality reduction, the grouping of clones across 24 different selected metrics and 27 different shape metrics utilized by the machine learning prediction model 204 is more easily visualized. FIG. 12 illustrates a comparison between a plot 1202 of PCA performed on image analysis using selected metrics (e.g., TER, gene expression, growth factor release, phagocytosis of photoreceptor outer segments, etc.) and a plot 1204 of PCA performed on image analysis using the machine learning prediction model 204. Between the two plots 1202 and 1204, the same separation tendency is seen for donors and clones. The first clone 1206 of Donor 2 is clearly separated farthest from the other clones, followed by the separation of the second clone 1208 of the same donor. Further, the clones 1210 of Donor 3 and the clones 1212 of Donor 4 are more strongly related to each other than the clones 1206 and 1208 of the other donors.
[0109] As shown in FIG. 13, this trend is consistent among the various machine learning techniques used by the machine learning prediction model 204. FIG. 13 illustrates a comparison of hierarchical clustering methods performed with and without using visual data according to an embodiment of the present disclosure. In FIG. 13, the first hierarchical clustering map 1302 represents the analysis of measurements using the selected metrics discussed above (without visual image data), and the second hierarchical map 1304 represents the visual image analysis performed by the machine learning prediction model 204. According to the illustrated embodiment, the first hierarchical clustering map 1302 and the second hierarchical clustering map 1304 are generated by applying multiscale bootstrap resampling to the hierarchical clustering of the analyzed data. Note that all numbers labeled by reference numeral 1306 represent the corrected probabilities of clustering, and the remaining numbers represent the uncorrected probabilities.
[0110] FIG. 13 illustrates that the analysis performed by the machine learning prediction model 204 not only matches the analysis performed using the selected metrics, but also provides valuable insights into the biology underlying what the analysis performed by the machine learning prediction model 204 can cause this grouping. The second hierarchical clustering map 1304 is consistent with the sequence analysis of iPSC-RPE cells for cancer genes (illustrated in FIG. 14), suggesting that the first clone 1206 of donor 2 had several mutated cancer genes during reprogramming.
[0111] FIG. 14 represents the sequence analysis of cancer genes of analyzed iPSC-RPE cells according to an embodiment of the present disclosure. FIG. 14 illustrates that nine clones from three different donors were tested. Only one of the tested clones, clone 1206, showed a mutation during reprogramming to iPSCs, which is represented by the highlighted region 1402.
[0112] In embodiments, certain software applications such as MATLAB® or Fiji may have limitations regarding the size of input images they can handle, so large input images can be split into overlapping tiles. Those tiles can be processed separately, for example, by performing segmentation using a deep neural network or by performing particle analysis i separately in each tile. Then, the image can be reconstructed using a suitable software application (e.g., C++), and the tiles can then be the C++ reconstructions from the processed individual tiles. The reconstructed image may not contain a visual indication that tiles were used.
[0113] In short, various embodiments of the present disclosure relate to a novel computational framework for creating lot and batch release criteria for clinical preparations of individual stem cell lines to determine the degree of similarity to previous lots or batches. Advantageously, this novel computational framework can be used for any stem cell type (e.g., but not limited to ESC, iPSC, MSC, NCS) that can be imaged using brightfield multispectral imaging, stem cell products from any source (e.g., cells derived from iPSC RPE), or any given cell line or genetically modified cell therapy product (e.g., chimeric antigen receptor (CAR) T cells). Advantages provided by the deep neural network model 212 and the machine learning prediction / classification model 204 contemplated by various embodiments of the present disclosure include the predictive ability of the models to automate culture conditions (e.g., but not limited to cell passage time, purity of iPSC or other stem cell types during culture, identification of whether healthy or unhealthy iPSC cell colonies, identification and potency of differentiated cells, identification and potency of drugs and toxins, etc.).
[0114] In other words, these techniques enable the convenient and efficient collection of a wide range of information, from the screening of new drugs to the study of the expression of new genes, the production of new diagnostic products, and the monitoring of cancer patients. This technology also enables the simultaneous analysis and isolation of specific cells. When used alone or in combination with state-of-the-art molecular techniques, this technology provides a useful method for correlating complex mechanisms, including the overall activity of living cells, with uniquely identifiable parameters. Furthermore, one of the important advantages provided by the disclosed computational framework is that the automated analysis performed by selected machine learning methods and state-of-the-art deep neural networks substantially eliminates human bias and errors.
Example
[0115] Experimental methods and procedures used A platform was developed using quantitative brightfield microscopy and a neural network for non-invasively predicting tissue function. In the experiment, it was determined whether the function of the tissue could be predicted from brightfield microscope images by using clinically graded iPSC-derived retinal pigment epithelium (iRPE) from age-related macular degeneration (AMD) patients and healthy donors as a model system.
[0116] The retinal pigment epithelium (RPE) is a monolayer of cells and is clinically interesting in studies related to the use of RPE for treating AMD. Moreover, the appearance of RPE cells within the monolayer is known to be extremely important for the function of the RPE, and in recent clinical trials, visual inspection of the RPE by skilled technicians has been used as a release criterion for biomanufacturing for transplantation. The appearance of RPE cells is roughly defined by the maturity of the junctional complexes between adjacent RPE cells and the characteristic pigmented appearance due to melanogenesis. The junctional complexes are related to the maturity and functionality of the tissue, including barrier functions (transepithelial resistance and potential (TER and TEP) measurements) and polarized secretion of growth factors (ELISA). Therefore, the appearance and function of the cells are correlated and can predict each other.
[0117] Due to the variability in images from optical microscopes, it is difficult to use those images for automated cell analysis and segmentation. Therefore, the platform developed in this study consists of two components. The first is QBAM, which uses an automated method to capture reproducible images between different microscopes. The second component is machine learning that predicts multicellular functions using the images (QBAM images) generated by QBAM. The machine learning methods are classified into the categories of deep neural networks (DNNs) and selected machine learning (SML). These methods were selected to provide speed, reproducibility, and accuracy along with a non-invasive automated method to assist in scaling the biomanufacturing process when cell therapies are translated from the laboratory to the clinic. Overview of the system and description of test cases
[0118] QBAM was developed to obtain reproducibility in brightfield imaging between different microscopes. QBAM converts pixels from relative intensity units to absorbance units, where absorbance is an absolute measure of light attenuation. To improve the reproducibility of imaging, QBAM calculates statistics on the image in real-time when the image is captured to ensure that the absorbance values measured at each pixel have a threshold confidence (configured at 95% confidence with a value of 10 milli-absorbance units (mAU) in this experiment). In this study, three different band filters were used for imaging, but this method can be adjusted to fit any number of wavelengths. QBAM imaging is run as a plugin for Micromanager (in the case of microscopes with available hardware) or the modular Python package (in the case of microscopes not supported by Micromanager), which is configured so that the user can obtain QBAM images with just a few button operations.
[0119] The analysis of QBAM images was selectively performed at the field of view (FOV) scale or the single-cell scale. For each of these scales, DNNs were used for different purposes. The DNNs at the FOV scale were designed to directly predict either the results of the function / maturity assay (via DNN-F) or whether two sets of QBAM images were from the same donor (via DNN-I). No image processing was performed before feeding the images into DNN-F or DNN-I.
[0120] Single-cell analysis started with a DNN that identified cell boundaries in QBAM images (via DNN-S). Next, the visual features of individual cells were extracted from the QBAM images using a web image processing pipeline (WIPP, the features that can be extracted are shown in Table 4 below). Then, using the extracted visual features, an SML algorithm was trained to predict various tissue characteristics, including the function, identity of the donor from which the cell originated, and developmental outliers (cells with an abnormal appearance). Next, the SML algorithm was used to identify the decisive cell features that contributed to the prediction of tissue characteristics. To demonstrate the effectiveness of this imaging and analysis method, proof-of-concept studies were conducted in iRPEs from the following donor types: healthy, oculocutaneous albinism disorder (OCA), and age-related macular degeneration (AMD). The iRPEs from healthy donors were imaged as they matured, while the AMD and OCA donors were imaged only once at the final time point when they had reached maturity.
[0121] iRPEs from five different OCA patients and two healthy donors were imaged using QBAM to determine the biological variability in OCA iRPEs and the sensitivity of this imaging method to the inherently low melanin levels. iRPEs from healthy donors were also evaluated for trans-epithelial resistance (TER) and polarized secretion of vascular endothelial growth factor (VEGF) in addition to weekly imaging. TER is a measure of RPE maturation and increases as tight junctions form between adjacent cells. Polarized secretion of VEGF is a measure of RPE function, with more VEGF secreted basolaterally compared to the apical side of the cell monolayer (VEGF-Ratio). Finally, eight iRPE clones were obtained from three AMD donors using a clinical-grade protocol. Here, clinical-grade refers to the generation of cells using xeno-free reagents and a cGMP-compliant manufacturing process. Once the AMD iRPEs reached maturity, QBAM imaging was performed on them. Accuracy, reproducibility, and sensitivity of QBAM
[0122] QBAM imaging was validated using a combination of a reference neutral density (ND) filter and biological samples. The ND filter with known absorbance values was used as a reference to validate QBAM imaging by comparing the absorbance measured in a UV-Vis spectrometer with the absorbance measured using QBAM imaging. The absorbance measured using QBAM imaging was strongly correlated with the absorbance measured by the spectrometer across the visible spectrum. To further validate this method, the reproducibility of QBAM of the ND filter was determined in three additional microscopes, each equipped with a different filter, objective lens, and light source. The root mean square error (RMSE) between all filters and microscopes was 66 milli-absorbance units (mAU), i.e., approximately 4.4% at the highest measured absorbance value.
[0123] Next, QBAM imaging was tested in gradually maturing primary iRPEs derived from two different healthy donors. A general trend of increasing average absorbance over time was observed. To determine how sensitive QBAM imaging is with respect to pigmentation in iRPE, iRPEs from five different patients with OCA (a disease known to have reduced pigmentation in iRPE) were imaged using QBAM. The OCA iRPEs were ranked to confirm the type of albinism (OCA1A or OCA2) and disease severity. The iRPEs of OCA1A had severe albinism and did not produce melanin (OCA8 and OCA26), and thus had the lowest image absorbance. OCA2 patients had phenotypes ranging from moderate (OCA103 and OCA9) to mild (OCA71), which correlated with the absorbance measurements performed by QBAM. Despite being iRPEs from OCA1A patients that produce low levels of pigment, the absorbance values were two-fold higher than the lowest sensitivity of QBAM (10 mAU). Collectively, these data demonstrate the accuracy, reproducibility, and sensitivity of QBAM imaging. Deep neural network prediction of iRPE function from absorbance images
[0124] To determine whether QBAM imaging affects cell maturation and can measure a wide range of variations in iRPE pigmentation, iRPEs from each healthy donor (Healthy-1, Healthy-2) were imaged. This was done using three culture conditions: (1) control iRPE (untreated), (2) iRPE treated with a known inducer of RPE maturation (aphidicolin), and (3) iRPE treated with a known inhibitor of RPE maturation (hedgehog pathway inhibitor-4, HPI4).
[0125] iRPE treated with control and afidicolin were found to mature as expected while increasing the absorbance of the images over 8 weeks of culture, while iRPE treated with HPI4 trended towards decreasing absorbance over time (referred to as Healthy-2 and Healthy-1). Higher expression of mature marker mRNAs and proteins was found in iRPE treated with control and afidicolin than in iRPE treated with HPI4. The baseline electrical responses (TEP and TER) and their changes to physiological treatments of 1 mM potassium (K+) or 100 μM adenosine triphosphate (ATP) were significantly greater in iRPE treated with afidicolin and significantly smaller in iRPE treated with HPI4 compared to the control. Furthermore, the maturation of iRPE was evident from the presence of natural-like dense apical protrusions. From this set of experiments, it can be concluded that (1) iRPE generated under clinical-grade conditions have a mature epithelial phenotype, (2) weekly QBAM imaging does not affect the maturation of iRPE, and (3) the difference in pigmentation between mature iRPE (control and afidicolin) and immature iRPE (HPI4) can be quantified using QBAM imaging.
[0126] Next, the ability to predict the function and phenotype of the iRPE monolayer from QBAM images of iRPE was evaluated. For iRPE from healthy donors, the mean value of the QBAM pixel values had little correlation with TER. However, the TER prediction by DNN-F was highly correlated with the actual TER values for the same samples, and these predictions had an RMSE of 70.6 Ω·cm 2 To incorporate this methodology into a biomanufacturing setting, a TER of 400 Ω·cm 2 was used to classify the iRPE monolayers as immature (<400 Ω·cm 2 ) or mature (>400 Ω·cm 2) was used as a stringent threshold for classification. False positives, false negatives, and TER values associated with immature iRPE were determined. Based on this TER threshold, for the classification of iRPE maturity, DNN-F was 94% accurate, with a sensitivity of 100% and a specificity of 90%. A similar trend was also observed for VEGF-Ratio, where the polarized release of VEGF was not well correlated with the QBAM average pixel value, but the DNN-F prediction was highly correlated with the measured value of VEGF-Ratio (R 2 = 0.89), and the RMSE of the VEGF-Ratio prediction was less than 1.0. The accuracy, sensitivity, and specificity of VEGF-Ratio were all 100% (where samples with VEGF-Ratio < 3.0 were considered immature). From these experiments, it can be concluded that (a) the QBAM images of live cells can be used to predict TER and VEGF-Ratio with high fidelity, and (b) QBAM imaging can be used as a non-invasive means of functional verification of cells instead of measuring TER and / or VEGF-Ratio. Extraction of single-cell features from live QBAM images of iRPE monolayers
[0127] DNN is known to have excellent predictive power compared to other machine learning algorithms, but it is difficult to determine which image features DNN uses to make predictions. To understand which cell image parameters of iRPE predict the function of a single layer, the features of images of individual iRPE cells in QBAM images were calculated and used to train an SML algorithm to predict iRPE function. A DNN was created (via DNN-S) to segment individual live iRPE cells in QBAM images. The DNN-S segmentation was verified by comparing the cell features calculated from 12,750 iRPEs with the same cell features calculated from manually corrected (ground truth) segmentation. When 44 different features were compared in the DNN-S vs. manually corrected segmentation (see Table 1 below), the difference between their feature histograms was shown to be 7.94% ± 4.42% (mean ± standard deviation), and they agreed well at the pixel level (F-2 = 0.71). Table 1: Manual measurement of 44 morphological features and quantification and distribution of % error from DNN-S segmentation
Table 1-1
Table 1-2
[0128] In Table 1, the % error indicates the absolute difference in numerical values at each value within the histogram. The two-sample Kolmogorov-Smirnov statistic (KSS), Kernel Maximum Mean Discrepancy (KMMD), and X2 value are three different ways to evaluate the difference in distributions. KSS is used in non-parametric distributions, KMMD has no assumptions about the initial distribution, and X2 is a chi-square test for normally distributed features.
[0129] QBAM imaging and live cell segmentation enable the measurement of features of hundreds of cell images and non-invasively track individual cells throughout the maturation of iRPE. Thus, using the trained DNN-S, QBAM images of untreated (control) or afidicolin- or HPI-4-treated primary iRPE (Healthy-1 and Healthy-2 donors) were segmented. Subsequently, previously published cell image features and intensity metrics known to correlate with RPE maturation and health were evaluated for significance. By observing the average number of neighbors each iRPE cell had in response to drug treatment, it was found that HPI4 had a significantly lower (p<0.001) average number of neighboring cells at all time points. Importantly, this approach enables unprecedented hierarchical granularity for the data; enabling measurements not only of the "bulk" tissue of the entire well, but also measurements according to the field of view or at the individual cell level. Clustering of treatments was performed based on two features known to be important for RPE maturation and health, cell area and average cell intensity. The minimum intensity of cells, a new metric for iRPE function identified using SML, is further described below. These results demonstrate that (a) the accuracy of DNN-S segmentation relative to manually drawn segmentation of individual cells and (b) differences between iRPE treated with different molecules can be explained using distinct cell features. Features of single cell images can predict iRPE maturation and function
[0130] Using five different SML methods (multilayer perceptron - MLP, linear support vector machine - L-SVM, random forest - RF, principal least squares regression - PLSR, and ridge regression - RR), cell features obtained from cell boundary segmentation of QBAM images were used to predict TER and VEGF-Ratio from iRPE of the Healthy-2 donor. The TER prediction of MLP had RMSE = 84.7 Ω·cm 2 and R2 The most accurate SML approach with = 0.94 was found. 400 Ω·cm 2 was used as the quality assurance or quality control (QA / QC) threshold to determine false positives and negatives. The MLP had 94% accuracy, 100% sensitivity, and 90% specificity. However, from the comparison of all algorithms, DNN-F was shown to be the most accurate predictor of TER, and the RMSE of the MLP was 14.1 Ω·cm 2 higher than that of DNN-F. For VEGF-Ratio prediction, the random forest (RF) was the best SML method, but the RMSE of DNN-F was 1.4 times lower.
[0131] The advantage of predicting the function of the iRPE monolayer from cell characteristics is that it is possible to determine the single-cell properties that suggest tissue-level function. For each SML method, the cell characteristics were ranked by importance. When comparing all SML models, there was similarity in the most important features for predicting TER or VEGF-Ratio regardless of which SML method was used. Interestingly, the important cell image features for predicting TER spread across cell intensity, texture, and shape. Of the 10 most important features, 3 were related to cell shape (shape), 3 were related to the pigment intensity within the cell (intensity), and 4 explained the distribution of pigment within the RPE (texture). Table 2 shows an excerpt of which metrics specifically represent these features, as well as their 95% confidence intervals for each time point and drug treatment. Table 2 Subsets of feature sets for modeling the Healthy-2 data, and their mean values, 95% confidence intervals (CI) and standard errors
Table 2
[0132] In summary, the above results indicate that using live cell segmentation, feature extraction, and SML can predict the TER and VEGF-Ratio of tissues at an accuracy level approaching that of DNNs for analyzing QBAM images. The advantage of the SML method compared to DNNs is that the SML model can identify distinct cell features that suggest the function of the iRPE monolayer, thereby enabling manufacturers to derive confidence intervals for the features of cell images against the release criteria for cell therapy. The accuracy of functional prediction is robust across iRPEs from multiple clinical-grade AMD patients
[0133] To determine the robustness across multiple donors and multiple preparations, DNN-F and SML were used to predict the TER and VEGF-Ratio of clinically graded iRPEs from three AMD patients in eight iPSC clones. The polarized phenotype of iRPE was confirmed by absorbance images of iRPE samples from each AMD donor and corresponding SEM images of iRPE apical protrusions. The maturation of the monolayer was evaluated by TER, VEGF-Ratio, and other assays. The average QBAM pixel values of iRPE were measured for fully mature AMD-iRPE. Similar to healthy donors, the average absorbance did not correlate well with TER (R 2 = 0.015) or VEGF-Ratio (R 2 = 0.50). However, the random forest (RF) model could predict TER to a similar level of accuracy (RMSE = 70.9 Ω·cm 2 , R 2 = 0.92) as healthy donors. DNN-F was also used to model TER, and the predicted and measured values were highly correlated (R 2 = 0.91).
[0134] To evaluate the robustness of the above approach, SML models were trained on different combinations of AMD-iRPE monolayers. A total of 18 unique training image subsets were formed, and each image subset included test data with images of one iRPE sample from each donor. The RMSE of the average TER was 86.9 Ω·cm 2 + / -14.3 Ω·cm 2 in all clone subsets (as shown in Table 5 below), indicating that the prediction error is similar when scaling to a larger donor subset (having 8 AMD-iRPE samples) compared to a single sample (Healthy-2). A 95% confidence interval was constructed from the measured and predicted values. iRPEs that fall outside this region may be considered "out-of-spec" in a biomanufacturing environment and may be recommended for further testing.
[0135] Finally, by evaluating the most important cell features for predicting the TER and VEGF-Ratio of AMD-iRPE monolayers in all donor / sample combinations, it was determined whether the features used to predict the function of AMD-iRPE were similar to those of Healthy-2. Interestingly, only four features, namely the Zernike n4-l0 polynomial (Shape 1), mass displacement (Intensity 2), and the third inverse difference moment at 135° (Texture 2) and 45° (Texture 1), overlapped between these two groups. Overall, the model derived from clinically graded iRPE images was able to predict the iRPE phenotype across multiple donors / samples and determine the features of live cell images common to multiple donors that predict the function of the iRPE monolayer. Furthermore, the difference in feature importance between Healthy-2 and AMD-iRPE suggests that both donor-specific features and features common to multiple donors may exist for predicting function. Classification of Developmental Outliers and Identities of iRPE Donors Using QBAM
[0136] By using QBAM images, it was determined whether any ontogenetic outliers were present based on the characteristics of cell images in eight clinical grade iRPE samples from three AMD patients. Ontogenetic outliers were defined as iRPE monolayers that differed from other iRPE monolayers based on the characteristics of the cell images, and further analysis could be warranted to determine whether that monolayer had developed appropriately. Principal component analysis was applied to the characteristics of cell images from QBAM images of AMD donor / sample preparations. iRPE from a given donor formed well-defined clusters with each other based on the characteristics of the cell images, except for AMD1 clone A (1A) and AMD3 clone A (3A), which were identified in the PCA dendrogram. Analysis of clone 1A iRPE showed 894 changes in the tumor exosome compared to the starting donor material. Clone 3A iRPE was found to have lower pigment levels than "sibling" iRPE. When analyzing the characteristics of their cell images, cell pigmentation and shape were found to be the two most dominant feature classes when discriminating these iRPE as outliers. A complete description of the features can be found in the online dataset.
[0137] For each iRPE monolayer, multiple SML models were used to predict the cell donor from the QBAM image. In addition, a new DNN (DNN-I) was developed to determine whether two different iRPE images were from the same donor. The SML algorithm took the features obtained from the QBAM image as input and the identification of the donor as output. In contrast, DNN-I took two QBAM images as input and classified them as being from the same or different donors. The SML approach could classify the donor identity of a new clone of a donor when trained on images of other clones from the same donor, but could not classify "new" donors that were not present in the training data. The DNN-I strategy for binary classification of two images as "same" or "not the same" gave the DNN-I the ability to classify "new" donors that were not used during training. The linear support vector machine (L-SVM) was found to have the highest accuracy among all the SML algorithms tested, with an accuracy of 76.4% (2.3× chance), a sensitivity of 64.6% and a specificity of 82.3%. In all donor / sample combinations, DNN-I had better performance, with an accuracy of 85.4% (2.6× chance), a sensitivity of 80.9% and a specificity of 86.8%. Interestingly, the important cell image features for differentiating AMD iRPE from each other were similar across different iRPE combinations and consisted of features different from those used to identify functional and developmental outliers of the tissue. The general differences among the top 50 features used in each application were as follows: Shape features were important for identifying clones as outliers (23 out of the top 50 features), texture features were important for donor classification of clones (25 out of the top 50 features), and shape and texture features were important for classifying iRPE function (40 out of the top 50 features). Considerations on Absorbance Imaging
[0138] Data input is crucial for the success of the analysis. Therefore, the image processing pipeline developed here starts with a rigorous absorbance imaging method using brightfield microscopy that is reproducible. The QBAM technology developed here can be implemented on any standard brightfield microscope and provides high confidence in image quality and reproducibility using a real-time automated statistically robust method. The advantage of using absorbance rather than raw pixel intensity is that absorbance is an absolute measure of light attenuation. Raw pixel intensity can vary depending on microscope setup and settings (e.g., non-uniform illumination, bulb intensity and spectrum, camera, etc.) that make image comparison difficult, even when the images are captured on the same microscope. By converting to absorbance values, many issues regarding image reproducibility are overcome. The combination of automating, converting pixel intensity to absorbance values, calculating absorbance confidence, and establishing microscope equilibration by benchmarking ensures the quality of image data captured using QBAM.
[0139] To ensure that its measurements can be used in multiple contexts, the robustness of QBAM was verified in three systems: (1) synthetic standards, (2) healthy biological samples, and (3) a drug-induced model of iRPE maturation. Analysis of QBAM images showed absorbance values that matched "known" synthetic standards and was able to assess pigment generation in both healthy and diseased RPE. This result also emphasizes the robustness of the measurements across multiple microscopes or imaging setups. The error between different microscope measurements of the same sample was within 4.4% of the signal, compared to average errors of 31% for controls against VEGF ELISA and 100% for TER measurements in the epithelium. This corresponds to a one- to two-order-of-magnitude reduction in variability when used in the context of a potential release assay for cell therapy products or in a drug screening platform.
[0140] QBAM is optimized for the measurement of the absorbance of cell types that absorb light. Thus, QBAM may be suitable for assays that use absorbance (e.g., counting live cells using trypan blue staining, histological staining, or analysis of light-absorbing biological specimens such as pigmented skin cells or dopaminergic neurons expressing neuromelanin). If transmittance values may be preferred (e.g., histological examination), the statistic may be modified to produce reproducible tissue section images. Also, since none of the calculated values are wavelength-specific, this method can be generalized to any multispectral modality. Finally, these methods are expected to be suitable for hyperspectral autofluorescence imaging that can identify cell boundaries and intracellular organization of non-pigmented cells. Considerations for Predicting iRPE Function
[0141] Neural networks and machine learning algorithms were trained using a variety of cell phenotypes by using healthy iRPE as well as drugs known to inhibit iRPE maturation (HPI4) and drugs known to promote iRPE maturation (aphidicolin). The presence of diverse phenotypes in the training set increased the robustness of the algorithm. Additionally, this method functioned not only as an endpoint assay for tissue health but also as a non-invasive tool for tracking tissue development over a long maturation period (about 35 days) for two different donors. Importantly, the accuracy of these algorithms in predicting both TER and VEGF-Ratio was close to the measurement uncertainty for both TER and VEGF.
[0142] As a cut-off ratio for biomanufacturing, 400 Ω·cm was selected for TER 2 , and 3 was selected for VEGF-Ratio. However, 200 Ω·cm 2 ~1000 Ω·cm 2When evaluating TER values in the range of or VEGF-Ratio in the range of 1 - 5, it is important to note that similar accuracy, sensitivity, and specificity were observed, and the threshold should be configured according to the manufacturer's specifications. In this study, a larger prediction error was observed for iRPE derived from Healthy-2 compared to the prediction error of iRPE derived from multiple AMD donors. Two reasons are hypothesized for this: (1) The iRPE of Healthy-2 had a wider range of both TER and VEGF-Ratio values than AMD-iRPE produced in a cGMP facility for the purpose of reproducibly producing a mature and healthy iRPE monolayer, due to including positive and negative controls. (2) Since the AMD-iRPE included all Healthy-2 data as well as training data from AMD-iRPE, the AMD-iRPE had a larger training dataset. In machine learning, generally, the more data there is, the more accurate the model obtained; however, in the execution of the platform in any application, it is necessary to specifically consider ensuring a wider range of conditions (i.e., more donors and positive / negative controls) for training than presented in this proof-of-concept study to ensure the robustness of the model.
[0143] As expected, deep learning had lower RMSE for the prediction of both TER and VEGF-Ratio compared to the SML approach. However, by using the SML approach, it was possible to discover important cell image features. Since two motivations were recognized for cell product manufacturers, regulatory authorities, clinicians and / or researchers, these two approaches were selected. Motivation 1: In manufacturing sites, clinical settings or high-throughput screening, time is often an important factor and a clear "continue / abort" or simple reading is desired. In these cases, the algorithm that provides the highest accuracy and is most robust to noise should be used, and deep learning is an excellent tool for this application. Motivation 2: In research, insights into the mechanisms underlying functions are often important. In these cases, a more interpretable method for determining the importance of cell image features is required. In the case of this motivation, the architecture underlying the SML approach is simple enough to be understandable and it is possible to obtain the importance of elements (here, cell image features) for the prediction of tissue function, so the SML approach is desirable.
[0144] By extracting features from QBAM images, hundreds of features can be obtained based on cell shape, intensity and texture at both the single cell level and larger cell populations. Many of these features are mathematical abstractions that have no meaningful connection to cell function. Therefore, SML models may be more interpretable than DNNs, but the features that make up these models may not be associated with the underlying biology. Nevertheless, for the purpose of manufacturing cells, what these features are and how they relate to the underlying biology; the ability to identify the 95% confidence interval and ensure that future batches / clones from donors fall within this range; is more important than justifying their use in SML models regardless of their relationship to the underlying biology. Consideration of the Classification and Clustering of Clinically Valid Donor-Derived iRPE
[0145] There is a highly important need to develop clinically compatible non-invasive assays to confirm the identity and quality of cell therapy products immediately prior to transplantation. PCA and cluster analysis can contribute to this unmet need. Using this approach or similar clustering methods, for the first time, the similarity between the future transplant and other technical replicas (or batches that have previously been successfully manufactured) can be non-invasively evaluated. Furthermore, the classification work performed using DNN-I or L-SVM can serve as a QA / QC process to detect manufacturing errors in cell implants and to match the identity of these implants with other replicas from the same donor. This is particularly important in facilities that manufacture thousands of autologous therapeutics and must confirm the identity of the administered product for each patient.
[0146] Two of the most important features for identifying developmental outliers were determined to be the standard deviation of the maximum intensity of iRPE (intensity 8) and the Zernike n5-l3 polynomial (shape 10). The deviation of its maximum intensity parameter is in good agreement with the absorbance results indicating that iRPE derived from clone A of AMD3 had lower absorbance than iRPE from other clones of AMD3. The Zernike polynomial was useful for detecting the shape of invasive cancer cells and classifying tumors, but this is hypothesized to be critical for detecting differences between clone A of AMD1 (which had 894 tumor exome mutations) and other AMD-iRPE strains. This cluster analysis may not be conclusive evidence that the occurrence of oncogene mutations can be revealed by QBAM imaging alone, but it is proposed that in the manufacturing environment of cell therapy, screening individual therapeutic replicas using cluster analysis may enable the determination of whether there are outliers that may require further scrutiny. Also, since this assay is non-invasive, this information can provide surgeons with additional confidence about the quality of the actual transplant delivered to the patient, which has not been possible until now.
[0147] In conclusion, the studies presented here show that by using QBAM imaging, the development of pigmentation in healthy and diseased iRPE can be non-invasively evaluated. DNN can analyze these images and accurately predict cell TER and VEGF-Ratio in 10 different iRPE preparations. Additionally, QBAM images contain sufficient information to enable DNN to accurately segment the RPE border of native RPE. Once segmented, hundreds of features can be calculated for each cell, and these features can be used to predict cell function, identify outlier samples, and confirm donor identity. All of this information can be obtained for tissue transplanted into patients in just a few minutes using an automated brightfield microscope, without the need for expert opinion from a clinician. Therefore, QBAM has the potential to be applied in a biomanufacturing environment to non-invasively test thousands of manufactured RPE units and qualify them for clinical use by technicians. Details of experimental models and subjects Origin and culture of iRPE cells Origin of human cells
[0148] All human research collection and processing was conducted under Protocol #11-E1-0245 approved by the Institutional Review Board at NIH. A total of 15 iRPE cell lines obtained from 10 different donors were used in this paper. Those iRPE lines were obtained from three types of patients: healthy patients, AMD patients, and OCA patients. iRPE from healthy patients (designated Healthy-1 and Healthy-2) were obtained from iPSC lines BEST4C and LORDY9, respectively. iRPE from AMD patients are referred to according to donor number and clone number. For example, AMD1A means cells derived from AMD donor #1 and clone A. Different clones for each donor are replicates, and each clone was completely reproduced from iPSC generation to iRPE differentiation. AMD clones were previously reported by Sharma et al., Patient-Specific Clinical-Grade iPS Cell-Derived Retinal Pigment Epithelium Patch Rescues Retinal Degeneration in Rodent and Pig Eyes, Sci. Transl. Med. Under Review (2018). A summary of the clone numbers for each donor is as follows: AMD1 had clones A and B, AMD2 had clones A, B, and C, and AMD3 had clones A, B, and C. iRPE obtained from OCA patients (also known as albinism patients) were from five different patients (each a single clone), designated OCA8, OCA26, OCA103, OCA9, and OCA71. Details of each donor's age and gender and information on the important sources are provided in Table 3. Table 3 Important sources
Table 3
[0149] Cell culture conditions and media All cells were cultured in a cell culture incubator at 37 °C and 5% CO2. Different cell culture media were used according to the developmental stage of the cells. The different cell culture media used are as follows:
[0150] Neural ectoderm induction medium (NEIM), which is DMEM / F-12 (Thermo Fisher, 11330-032), KOSR (CTS KnockOut SR XenoFree Kit, Thermo Fisher, A1099201), supplemented with 1% (mass / volume) N-2 (Thermo Fisher, A13707-01), 1× B-27 (Thermo Fisher, 17504-044), 10 μmol / L of LDN-193189 (Stemgent, 04-0074-10), 10 nmol / L of SB431452 (R&D Systems, 1614), 0.5 μmol / L of CKI-7 dihydrochloride (Sigma Aldrich, C0742-5mg) and 1 ng / ml of IGF-1 (R&D Systems, 291-GMP-5.5ug).
[0151] RPE induction medium (RPEIM), which is DMEM / F-12 (Thermo Fisher, 11330-032), KOSR (Thermo Fisher, A1099201), supplemented with 1% N-2 (Thermo Fisher, A13707-01), 1× B-27 (Thermo Fisher, 17504-044), 100 μmol / L of LDN-193189 (Stemgent, 04-0074-10), 100 nmol / L of SB431452 (R&D Systems, 1614), 5 μmol / L of CKI-7 dihydrochloride (Sigma Aldrich, C0742-5mg), 10 ng / ml of IGF-1 (R&D Systems, 291-GMP-5.5ug) and 1 μmol / L of PD0325901 (Stemgent 04-0006).
[0152] RPE commitment medium (RPECM), which is DMEM / F-12 (), KOSR (Thermo Fisher, A1099201) supplemented with 1% N-2 (Thermo Fisher, A13707-01), 1× B-27 (Thermo Fisher, 17504-044), 10 mmol / L nicotinamide (Sigma Aldrich, PHR1033-1G), and 100 ng / ml activin A (R&D Systems, AFL338).
[0153] RPE growth medium (RPEGM), which is MEMα (Thermo Fisher, 12571-063) supplemented with 1% N-2 (Thermo Fisher, A13707-01), 1% (mass / volume) GlutaMAX Supplement (Thermo Fisher, 35050-061), 1% (mass / volume) non-essential amino acids (MEM non-essential amino acid solution (100×), Thermo Fisher, 11140-050), 0.25 mg / ml taurine (Sigma Aldrich, PHR1109-1G), 20 ng / ml hydrocortisone (50 μmol / L solution, Sigma Aldrich, H-6909), 13 pg / ml triiodothyronine (Sigma Aldrich, T5516), and 5% (volume / volume) FBS (fetal bovine serum, GE Healthcare / Hyclone, SH30071.03).
[0154] RPE maturation medium (RPEMM) is MEMα (Thermo Fisher, 12571-063) supplemented with 1% (mass / volume) N-2 (Thermo Fisher, A13707-01), 1% (mass / volume) glutamine (Thermo Fisher, 35050-061), 1% (mass / volume) non-essential amino acids (Thermo Fisher, 11140-050), 0.25 mg / ml taurine (Sigma Aldrich, PHR1109-1G), 10 ng / ml hydrocortisone (Sigma Aldrich, H-6909), 13 pg / ml triiodothyronine (Sigma Aldrich, T5516), 5% (volume / volume) FBS (GE Healthcare / Hyclone, SH30071.03), and 50 μM PGE2 (prostaglandin E2, R&D Systems, 2296). Transformation and Differentiation of Human iRPE
[0155] All iRPE were generated using a previously outlined clinical-grade protocol (Sharma et al., 2018). Briefly, a previously published protocol (Mack et al., Generation of Induced Pluripotent Stem Cells from CD34+ Cells across Blood Drawn from Multiple Donors with Non-Integrating iPSC clones were obtained from CD34+ PBMCs using Episomal Vectors, W.B. (2011). Subsequently, the iPSC cells were seeded onto tissue culture plates (Thermo Sci., 140675) or T75 flasks (Corning, 430641U) containing E8 medium (Essential 8 Medium, ThermoFisher, A1517001) coated with 5 μg / ml vitronectin (ThermoFisher, A14701S). After 2 days, the cell medium was replaced with RPEIM for 10 days and then further replaced with RPECM for another 10 days. On day 22, the cell medium was switched to RPEGM. On day 27, the cells were detached using Versen solution (0.2 g of EDTA (ethylenediaminetetraacetic acid) per liter of phosphate-buffered saline, Thermo Fisher, 15040-066) and re-seeded onto a new culture plate (Thermo Sci., 140675) or T75 flask (Corning, 430641U) containing RPEGM. On day 42, the cells were detached using CTS TrypLE Select Enzyme (Thermo Fisher, A12859-01) and re-seeded at 500,000 cells / ml onto a biodegradable nanofiber scaffold [Stellenbosch Nanofiber Company, FiberScaff-RPE TM 3D cell culture scaffold] (AMD Co., Ltd.) or a 12-well 0.4 μm polycarbonate Transwell plate (Corning, 3401) (Healthy Co., Ltd. and OCA Co., Ltd.) and cultured using RPEMM. The RPE seeded on day 42 was considered as time 0 iRPE, and all timings for drug treatment, imaging, sample collection for assays, etc. in the experiments outlined in this document were counted from this day. Details of the method Quantitative brightfield absorbance microscopy Basic principles
[0156] The basic principle underlying QBAM imaging is the absorbance measurement value, which is the absolute measurement of light attenuation as described in Equation 1.
Number
[0157] A is the absorbance value at wavelength A, I(λ) is the intensity of light passing through the sample, and I0(λ) is the intensity of light in the absence of the sample. One reason why absorbance is interesting in this study is that Lambert-Beer's law establishes the relationship between absorbance and chemical concentration as shown in Equation 2: [Number]
[0158] C is the chemical concentration in the sample, ε is the molar attenuation coefficient, and ι is the path length of the light beam. In the case of retinal pigment epithelial cells (RPE), as they mature, melanin is produced by healthy RPE. Therefore, a doubling of absorbance may suggest a doubling of the amount of melanin in the cells. By converting the pixel values in the RPE image to absorbance values, the image becomes a melanin concentration map that can be tracked over time and non-invasively. In addition to the relationship with melanin concentration, there were several advantages to using absorbance imaging, the most important of which was reproducibility. To calculate absorbance values, an internal standard is required to allow comparison of values between microscopes with different configurations. Using QBAM imaging, each pixel value was converted to an absorbance value by dividing the intensity (I(λ)) of each pixel in the sample image by the corresponding pixel (I0(λ)) in the image captured in the absence of the sample. As with most measurements, there are several factors to consider when making measurements to ensure reproducibility. QBAM imaging reduces some of the sources of error through various procedures including benchmarking, internal calibration, and real-time statistics. Calculation of Pixel-Level Absorbance
[0159] In the case of QBAM imaging, three different images were required to calculate absorbance at each pixel: i) an image (IDark (λ)), ii) the image (I Bright (λ)) captured with the optical shutter open and the field of view not containing the sample at the exposure time ε, and iii) the image (I(λ)) of the sample. Regarding Equation 1, (I Bright (λ)) was the blank reference image (I0(λ)), but as described in Equation 3, additional terms needed to be included for (I Dark (λ)) to account for the camera bias and readout current:
Equation
[0160] The subscripts i and j indicate the positions of the pixels in the i-th row and j-th column of each image. Note that the calculation of the transmittance (the term inside this log function) is the pre-background correction method reported by Young, 2001 (Young, Shading correction: compensation for illumination and sensor inhomogeneities. Curr. Protoc. Cytom. Chapter 2, Unit 2.11, (2001)). This means that the calculation of the pixel-level absorbance essentially corrects for non-uniform illumination and is one factor that makes QBAM imaging robust across microscope setups. Calculation of Pixel-Level Confidence Intervals
[0161] According to Equation 3, there were three different measured values taken to calculate the absorbance, each with its own potential source of error. To improve the reproducibility of the pixel-level absorbance measurements, the QBAM imaging method uses statistics to calculate a confidence interval for the absorbance value at each pixel. The goal of these statistical methods was to capture enough image data to ensure that the absorbance value at each pixel had a 95% confidence interval of 0.01 absorbance units (10 mAU). This means that the lower end of the dynamic range of QBAM is 10 mAU.
[0162] The standard deviation of the absorbance value was calculated according to Equation 4:
Number
[0163] In Equation 4, σ A(λ)i,j is the standard deviation of the absorbance value at position (i, j). σ I(λ)i,j and σ IBright(λ)i,j are, respectively, the standard deviation of the pixel intensity value I(λ) for the pixel at position (i, j) in the sample image i,j and the standard deviation of the pixel intensity value I Bright(λ)i,j for the pixel at position (i, j) in the bright reference image. Equation 4 was derived using the propagation of uncertainty, but two things should be noted about this derivation. First, it was assumed that I(λ) i,j and I Bright (λ) i,j are not correlated. This has not been verified experimentally, but if these variables are correlated, additional terms are subtracted from Equation 4, which means that the standard deviation of the absorbance has been overestimated by this equation. Second, since the standard deviation of the dark reference image accounted for less than 1% of the standard deviation of the absorbance, the error introduced by the dark reference image I Dark (λ) was ignored.
[0164] To ensure the reproducibility of the absorbance measurement, a criterion was set for the calculation of the pixel such that the absorbance value of the pixel has a 95% confidence interval of 0.01 absorbance units. It is assumed that it follows a normal distribution as shown in Equation 5.
Number
[0165] where n is the number of images captured for I(λ) i,j and n is simply the number of images captured of the sample, and I BrightNote that since it does not include the number of captured images, the right side of Equation 5 represents an overestimation of the confidence interval. Reduction of Chromatic Aberration by Color Filters
[0166] When comparing images between different transmission optical microscopes, one reason the images may appear different is the light spectrum. Everything related to how the light spectrum is generated, how it is manipulated using optical components, and how it is collected in the microscope can vary from microscope to microscope. For example, different light sources emit different light spectra that can change with age and temperature, different objective lenses correct chromatic aberration differently, and different cameras have different sensitivities to light of different wavelengths. Absorbance measurements can reduce issues related to the spectral emission of the light source and the spectral sensitivity of the camera for a single microscope setup, but color filters need to be reproducible between imaging sessions and between microscopes. In the case of this study, images of fixed RPE (AMD and OCA cells) were captured on a Zeiss AxioImager M2 microscope using three different color filters: 405 nm (ET405 / 10x, Chroma, Bellows Falls, VT), 548 nm (ET548 / 10x, Chroma, Bellows Falls, VT), and 640 nm (ET640 / 20x, Chroma, Bellows Falls, VT). Images of live RPE (Healthy-2 cells) were captured on a Zeiss AxioObserver Z1 using three different color filters: 461 nm (FF01-461 / 5-25, Semrock, Rochester, NY), 541 nm (FF01-541 / 3-25, Semrock, Rochester, NY), and 671 nm (FF01-671 / 3-25, Semrock, Rochester, NY). Microscope Benchmarking and Capturing of Blank Images
[0167] The purpose of benchmarking is to determine whether the pixel intensity of a microscopic image responds directly to changes in illumination. In the case of QBAM imaging, doubling the exposure time should double the pixel intensity. Therefore, this determination is achieved by finding out how the pixel intensity changes when the camera's exposure time is varied.
[0168] The benchmarking protocol calculates the following quantities: pixel intensity as a function of exposure time, variance of pixel intensity as a function of average pixel intensity, optimal exposure time, and minimum allowable pixel intensity. These quantities and their calculations are described in detail below. Since the characteristics of the image can vary with wavelength, this benchmarking protocol is used for each image filter.
[0169] The first step of the benchmarking protocol involves capturing images with the camera's shutter closed and open at various exposure times. Before capturing the images, the user defined the number n of images to be captured at each exposure time. Then, the camera's shutter was closed, and n images were captured at 1 ms, 2 ms, 4 ms, etc. up to 1024 ms, and all the images were saved for quality control. Then, the camera's shutter was opened, and n images were captured at 1 ms, 2 ms, and 4 ms. For each set of images captured at one exposure time, the average and standard deviation of the pixel intensity were calculated for each pixel. Then, the exposure time was doubled, n images were captured, the average and standard deviation were calculated for each pixel, and this process was repeated until the image was overexposed. Overexposure at a particular exposure time was determined as follows.
Equation
[0170] where σ 2 I(λ,ε)i,j is the wavelength λ and the exposure time 2 εThe variance of the intensity at pixel (i,j) captured at time ms, where I*J is the total number of pixels in the image. Since overexposure is defined by reaching the maximum pixel intensity value of the camera, the standard deviation of overexposed pixels is 0. Therefore, for exposure time 2 ε when the captured image at ε-1 has less variation than the image captured at the previous exposure time 2
Equation
[0171] where σ λ (p) is the standard deviation of the pixel intensity with respect to the pixel intensity p at wavelength λ, and α and β are parameters calculated from the linear regression of the average pixel intensity with respect to the standard deviation in the images captured at 1ms, 2ms, and 4ms.
[0172] Next, by performing a linear regression, we determined how the pixel intensity changes with the exposure time. The camera exposure times that resulted in overexposed images were excluded from the linear regression, and if an image captured at a certain exposure time had an error of more than 5% with respect to Equation 7, that image was classified as overexposed. According to the linear regression of the pixel intensity with respect to the exposure time, the ideal exposure time was calculated in Equation 8 as follows:
Equation
[0173] where ε λ is the ideal exposure time, and p Dis the maximum camera pixel bit depth, and σ λ (p D ) is the estimated standard deviation (from Equation 7), and a and b were the slope and intercept of the linear regression of pixel intensity against exposure time. The ideal exposure time is the time at which the mean exposure time should be 3 standard deviations lower than the maximum possible pixel intensity, thereby minimizing the number of individually overexposed pixels.
[0174] Once the ideal exposure time is determined, I Bright is captured to obtain, and the number of images to be averaged is calculated using Equation 9 with the estimated standard deviation function of Equation 7:
Equation
[0175] where n Bright is the number of images required to have a 95% confidence interval E for the pixel intensity p ε in the ideal exposure. In this study, E was set to 2 pixel intensity units, which corresponded to an error of approximately 0.05% for the 12-bit camera used in this study. By setting E = 2, an approximation of n Bright ≈ σ λ (p ε ) 2 becomes possible, but this approximation should not be used for cameras with different bit depths.
[0176] The minimum pixel intensity ( p min) for accurately calculating absorbance was determined. Absorbance is a function of the logarithmic scale, and pixel intensities lower than I Bright can have a much larger error than higher pixel intensities. During live cell imaging, if any pixel in an image has a value smaller than the minimum pixel intensity, there is a high probability that the absorbance value of that pixel does not have a 95% CI less than 0.01. The following Equation 10 estimates the minimum pixel intensity value that satisfies Equation 5 using Equation 7 and Equation 4:
Number
[0177] Adjust the method of capturing an image to satisfy Equation 5 by using the minimum pixel intensity value during live cell imaging. Evaluation of Microscope Equilibrium
[0178] Equilibrate the microscope used for QBAM imaging so that it can capture images with little variation in the same field of view. Due to multiple factors, how the images taken in the same field of view can change from one image capture to the next (e.g., the change in the intensity of a light bulb as it gets hot, or the change in the sensitivity of a camera as it gets hot by being used) can vary. When images are repeatedly captured, the microscope should reach an equilibrium state where the consecutive images have minimal variation between them. By using the benchmarking method described above, an equilibrium metric was created that is measured when the microscope reaches equilibrium. That equilibrium metric is shown in Equation 11:
Number
[0179] In the formula, E q is the equilibrium metric, a t is the slope obtained from the linear regression of pixel intensity with respect to the exposure time described in the benchmarking section,
Number
[0180] After ensuring that the microscope was in equilibrium, the microscope was benchmarked, and reference images were captured (I Bright and I Dark ), and then the samples were imaged. For each field of view of the sample, n images were captured and averaged. The bright reference image was always captured with the microscope focused on the same medium in which the sample was prepared. For fixed cell samples mounted on microscope slides, the bright reference image was captured with the microscope focused on the blank section of the microscope glass. For live cell imaging, the bright reference image was captured in a well containing the same volume of the medium in which the cells were cultured. If the average of any pixel values was less than the minimum pixel value p min calculated during benchmarking, the ideal exposure time was doubled, and an additional n images were captured and averaged. This process of doubling the exposure time and capturing additional images was repeated until all pixels in the averaged image were pmin This was repeated until it became larger. All the images captured at each exposure time were saved and used when converting from pixel values to absorbance values to enhance the reliability of the measurement.
[0181] To calculate the absorbance for each field of view, Equation 3 was used together with the average value of the n images used for I(λ). i,j For I(λ), the images captured at longer exposure times were used, and the absorbance at that pixel was calculated by dividing the pixel intensity by the doubling rate of the exposure time. Since the pixel intensity follows a Poisson distribution, dividing the pixel intensity by the doubling rate of the exposure time makes the standard deviation smaller by the same amount and the confidence interval narrower. i,j <p min In the case of, for I(λ), the images captured at longer exposure times were used, and the absorbance at that pixel was calculated by dividing the pixel intensity by the doubling rate of the exposure time. Since the pixel intensity follows a Poisson distribution, dividing the pixel intensity by the doubling rate of the exposure time makes the standard deviation smaller by the same amount and the confidence interval narrower. Culture, assay, and imaging of iRPE Immunostaining
[0182] The preparation of the tissue for immunohistochemistry was performed by placing 4% (mass / volume) paraformaldehyde (Electron Microscopy Science, 157-4-100) in the wells for 20 minutes. The immunohistochemistry blocking solution (IBS) consisted of 500 ml of 1× DPBS (Dulbecco's phosphate buffered saline, Life Technologies, 14190250), 5% (mass / volume) bovine serum albumin (Sigma Aldrich, A3311), 0.5% (mass / volume) Triton® X-100 (Sigma Aldrich, X100-100ML) and 0.5% (mass / volume) TWEEN® 20 (Sigma Aldrich, P2287-100ML). The fixed cells were washed three times with IBS and permeabilized with IBS at room temperature for 2 hours. The cells were then stained with the following primary antibodies for 1 hour at room temperature: RPE65 (anti-RPE65 monoclonal antibody [1:300, Abcam, ab13826), ezrin (monoclonal anti-ezrin antibody, 1:100, Sigma Aldrich, E8897) or GT335 (anti-polyglutamylation modification monoclonal antibody, 1:1000, Adipogen, AG-20B-0020). After primary staining, the samples were washed with IBS and the goat anti-mouse IgG Alexa Fluor 594 secondary antibody (1:300, Thermo Fisher, A-11032) was added and incubated overnight at 4°C. All antibodies were diluted in the IBS solution. After overnight incubation, the samples were washed with IBS and the anti-ZO-1 Alexa Fluor 488 mouse monoclonal antibody (1:200, Thermo Fisher, 339188) was added and incubated for 1 hour at room temperature. In addition, the nuclei were stained with Hoechst 33342 dye (1:1000, Thermo Fisher, H3570) for 15 minutes at room temperature. After staining, the cells were washed with D-PBS (Dulbecco's phosphate buffered saline, Life Technologies, 14190250) and mounted on slides. All images were captured using a Zeiss AxioImager M2 microscope or a Zeiss Axio Scan Z1 slide scanner.Z-stacks were acquired along the z-axis at 1.5-μm intervals over 50 μm, and maximum intensity projection was used for all analyses. ELISA quantification of VEGF
[0183] Vascular endothelial growth factor (VEGF) was measured from the supernatants collected weekly from each well. At least 5 well replicates were selected for each treatment at each time point, and 3 technical replicates per well were measured. ELISA was performed using a singleplex Luminex® multiplex ELISA according to the manufacturer's protocol (R&D Systems). Electrophysiological measurements
[0184] Transepithelial electrical resistance measurements used for the prediction of iRPE monolayers were measured using an EVOM2 and an EndOhm chamber (World Precision Instruments, EVOM2 and ENDOHM-12, respectively). Each strain was measured for its resistance at 12 wells per treatment per week. Measurements of intracellular transepithelial potential and resistance at 1 mM potassium (Sigma Aldrich, P5405-1KG) and 100 μM ATP (Sigma Aldrich, A9187-500MG) were performed as previously reported (Miyagishima et al., In Pursuit of Authenticity: Induced Pluripotent Stem Cell-Derived Retinal Pigment Epithelium for Clinical Measurements were performed as described in Applications, STEM CELLS Transl. Med. 5, 1562 - 1574 (2016). Briefly, RPE monolayer cultures derived from iPSC lines were mounted in a modified "Ussing" chamber as previously reported (Maminishkis et al., The P2Y2 Receptor Agonist INS37217 Stimulates RPE Fluid Transport In Vitro and Retinal Reattachment, Rat. Invest. Ophthalmol. Vis. Sci. 43, 3555 - 3566 (2002); Peterson et al., Extracellular ATP Activates Calcium Signaling, Ion, and Fluid Transport in Retinal Pigment Epithelium, J. Neurosci. 17, 2324 - 2337 (1997)). The transepithelial potential (TEP) was measured using calomel electrodes in series with a Ringer's solution and an agar bridge. The signal from an intracellular microelectrode was measured as the basolateral membrane potential (V b ) with respect to the basal bath, and the apical membrane potential (V a ) was calculated using the equation: V a = V b - TEP. A current pulse of 2 - 4 mA was passed through the tissue, and the changes in the resulting TEP, V a and V b were measured to obtain the total transepithelial resistance (R t ) and the ratio of the apical membrane resistance to the basolateral membrane resistance (R A / R B ). Gene expression of iRPE
[0185] Total RNA was isolated using NucleoSpin RNA (Machery-Nagel, 740955) according to the manufacturer's protocol. RNA was quantified using an ND-1000 spectrophotometer (Nanodrop Technologies) and the manufacturer's protocol. cDNA synthesis was performed using a Script cDNA Synthesis Kit (Bio-Rad, 1708891) and the protocol provided by the manufacturer. A custom gene array containing primers for the following genes: MITF, PAX6, BEST1, CLDN19, PMEL, TYR, OCA2, RPE65, RLBP1 and the housekeeping genes RPL13A and B2M was purchased from Bio-rad. Sybr green-based QPCR was performed on a ViiA 7 Real-Time PCR System (Thermo Fisher Scientific) using the EXPRESS One-Step SYBR GreenER Kit according to the manufacturer's protocol with premixed ROX (Thermo Fisher, 1179001K). Each sample was assayed in at least two independent wells in triplicate technical replicates (2×HPI4, 3×affidicolin and 5×control). Wells of HPI4 had to be excluded because of extremely low RNA extraction amounts and insufficient amounts of cDNA amplified for measurement. The mean CT of the housekeeping gene for each well was used to calculate the plotted -ΔCT values. Data were analyzed using R software (R Development Core Team, 2011). Electron microscopy of iRPE
[0186] The mature iRPE monolayer was fixed overnight with 10% (mass / volume) glutaraldehyde (Electron Microscopy Sciences, 16120). The samples were then washed three times with D-PBS and immersed in 25% (mass / volume) ethanol for 10 minutes. The samples were then removed and successively placed in solutions of 50%, 75%, 90%, 95% and 100% ethanol for 10 minutes each. After immersion in 100% ethanol, the samples were removed, placed in a critical point dryer (Leica EM CPD300) and processed for erythrocytes according to the manufacturer's protocol. After CPD treatment, the samples were gold sputter-coated with 10 nm gold and imaged using a Zeiss EVO 10 scanning electron microscope at a voltage of 10 kV and an ampere number of 10 μA. Drug treatment and imaging
[0187] After differentiation, iRPEs from Healthy-1 and Healthy-2 donors were seeded at a density of 500,000 cells / ml in 0.5 ml into 12-well 0.4-μm polycarbonate transwell plates (Corning, 3401). By using only 6 out of the 12 wells, the time spent outside the incubator during imaging was minimized. QBAM imaging and drug treatment were started 2 or 1 week after seeding of iRPEs from Healthy-1 and Healthy-2, respectively. For each culture plate, 2 out of 6 wells were treated with 3.0 μmol / L aphidicolin (Sigma Aldrich, A0781-10MG) or 30.0 μM HPI4 (hedgehog pathway inhibitor 4, Sigma Aldrich, H4541-25MG), or treated without additives to RPEMM medium. RPEMM was changed three times a week, and imaging was performed on the same day every week immediately after medium change. Since the medium at room temperature minimized condensation, the culture medium in the wells was exchanged with RPEMM at room temperature (containing drugs if necessary) immediately before imaging to prevent the lid of the culture plate from fogging during imaging. To capture bright reference images for QBAM imaging, RPEMM without drugs was added to wells without cells. For each well, overlapping images of a 4×3 grid were captured using a Zeiss AxioImager M2 microscope with a 10× objective lens for image stitching (10-15% image overlap). For Healthy-2 and Healthy-1, a total of 6 plates were prepared, and since 6 wells were occupied in each plate, a total of 36 wells were imaged and analyzed. iRPEs from AMD samples were previously reported by Sharma et al. (Sharma et al., 2018), and iRPEs from OCA donors were cultured as described above, but were imaged only after they matured and were fixed with PFA. A total of 5 replicates for OCA8, 6 replicates for OCA9, 4 replicates for OCA26, 8 replicates for OCA71, and 4 replicates for OCA103 were imaged.In the case of AMD data, one iteration of AMD1 clone A, five iterations of AMD1 clone B, one iteration of AMD2 clone A, five iterations of AMD2 clone B, one iteration of AMD2 clone C, one iteration of AMD3 clone A, five iterations of AMD3 clone B, and one iteration of AMD3 clone C were imaged.
[0188] In the case of iRPE imaged from the Healthy-2 donor, images were captured using three color filters, while Healthy-1 was imaged using a different filter set. The filter used to image Healthy-1 was a filter that allowed light with a wider bandwidth to pass through, which had the advantage of shorter exposure times and, as a result, the advantage that the imaging cells spent less time outside the incubator. To ensure that iRPE maturation was not adversely affected by imaging, Healthy-1 was first imaged using a wide-band filter. Once it was proven that the iRPE maturation of Healthy-1 was not adversely affected by QBAM imaging, a filter with a narrow light bandwidth was used for imaging Healthy-2 cells. All of the above images were acquired using a Zeiss AxioImager M2 microscope. Deep neural network prediction of the assay Network architecture and training
[0189] A deep convolutional neural network (DNN-F) was designed to predict transepithelial resistance (TER) and VEGF-Ratio. The basic structure of this network consisted of a series of preliminary convolutional layers followed by inception layers deeper in the network (Szegedy et al., Going Deeper with Convolutions, ArXiv14094842 Cs (2014)). This network takes a 1024×1024×3 image as input and generates two values: TER and VEGF-Ratio values as output. This network was first trained to predict TER values, and then a VEGF-Ratio prediction layer was added at the end of this network. The 3-channel image used as input was composed of three QBAM images captured at different wavelengths (488nm, 561nm, and 633nm). This network had approximately 11 million parameters.
[0190] Before training, the measured values of cells were scaled and weighted to improve the prediction ability of DNN-F. By dividing the TER measurement values by 1000 and the VEGF-Ratio measurement values by 10, those values were roughly scaled to the range of 0 to 1 for training. Once training started, equal weights were given to all images, but in the later stage of training, weights were given to each image based on the TER value associated with that image. Weights were generated by binning the TER values into bins with a range of 100 Ohm, where the smallest bin included TER values in the range of 0 to 135 Ohm. Then, the weights were assigned by dividing the total number of TER measurements by the number of measurements in each bin. As training progressed, the weights were gradually decreased until all weights had a value of 1.
[0191] For training, mean squared error (MSE) regression was used as the objective function. Standard stochastic gradient descent was used with a constant learning rate. No image preprocessing other than random cropping of the images was performed before feeding the images into the network. Each image was 1040×1388 pixels, and the images in the training data were randomly cropped just before feeding them into the network, while the test data were cropped at the same position each time. Test data and training data were created by assigning one culture plate as the test data and using the remaining plates as training data (2 replicates per treatment per plate; aphidicolin, HPI4, control). The best network was determined by finding the network that used the lowest MSE in the test set. Fluorescent image segmentation
[0192] A deep convolutional neural network was designed to segment RPE that was fluorescently labeled for tight junction protein (ZO-1), thereby highlighting cell boundaries and enabling accurate cell segmentation. Since the cell boundaries in the QBAM images have lower contrast than the fluorescent images of RPE and it is difficult to segment the cell boundaries manually, the above objective was to obtain a highly accurate segmentation method and generate the correct cell boundary labels for QBAM. This approach consisted of 1) training a DNN (DNN-Z) to segment cell boundaries in the ZO-1 fluorescence images using the corresponding images (where the cell boundaries were drawn by an expert technician), 2) collecting QBAM images and fluorescence images of RPE fluorescently stained for ZO-1, 3) using DNN-Z to segment cell boundaries using the ZO-1 fluorescence images, and 4) training a new DNN (DNN-S) to segment cells in the QBAM images using the ZO-1 segmentation. Human and mouse data
[0193] Images of retinal pigment epithelial (RPE) cells labeled against the tight junction protein zonula occluden-1 (ZO-1) were collected from various human donors under various conditions. Human RPE was induced from induced pluripotent stem cells (iRPE), and the iRPE was obtained from healthy human donors or human donors having one of two different disease phenotypes: oculocutaneous albinism (OCA) or age-related macular degeneration (AMD). The images of iRPE were derived from a total of 10 different human donors. All iRPE were matured for at least 6 weeks, fixed with paraformaldehyde, labeled against ZO-1, mounted on microscope slides, and then imaged. The images were captured at magnifications in the range of 10× to 40× on three different microscopes. To that human data, images of the entire adult mouse retina labeled for cell boundaries and labeled with ZO-1, provided courtesy of the laboratory of John Nickerson, were further added. Reference segmentation images and post-processing
[0194] Among many segmentation methods, the FogBank algorithm (Chalfoun et al., FogBank: a single cell segmentation across multiple cell lines and image modalities, BMC Bioinformatics 15, 431 (2014)) was selected to segment each of the fluorescence-labeled images. The FogBank algorithm uses thresholding obtained from intensity distributions in combination with a geodesic distance map of the contour to define the region of RPE cells. The results of FogBank segmentation were reviewed and corrected by human subjects to obtain the correct segmentation data. After creating that correct data, the images were divided into 256×256 tiles to obtain 4,064 images of size 256×256 pixels. Architecture and training of a deep convolutional neural network for RPE segmentation
[0195] A teacher-guided DNN-based segmentation algorithm was designed in MATLAB (R2017a) using MatConvNet (Vedaldi et al., MatConvNet - Convolutional Neural Networks for MATLAB. ArXiv14124564 Cs(2014)), an open-source machine learning framework. Using the same architecture, fluorescent images labeled for ZO-1 (cell tight junction staining) were segmented, and cells in QBAM images were also segmented. The basic layer structure is the inception layer used in GoogLeNet (Szegedy et al., 2014), and its higher-order structure follows the U-Net architecture (Ronneberger et al., U-Net: Convolutional Networks for Biomedical Image Segmentation. In Medical Image Computing and Computer-Assisted Intervention - MICCAI, (Springer, Cham), pp. 234 - 241 (2015)). This network takes an image of 256×256 pixels as input and outputs an image of 200×200 pixels, where the cell boundary has a positive pixel value and the cell body has a negative pixel value. This network had more than 15 million parameters.
[0196] To train the above network, stochastic gradient descent was used with ADADELTA optimization (ε = 10 -6 and ρ = 0.9). Batch normalization was used with 10 images in each batch. The F1 score was used to determine the best network model, and correctly labeled cell boundaries were considered true positives. A dropout layer was used with a 50% dropout rate before the last convolution used to classify pixels. This objective function was a modified logistic log-loss function as shown in Equation 12:
Equation
[0197] Where x is the predicted pixel class, c is the actual pixel class, c = +1 represents the cell boundary, c = -1 represents the cell body, and B = max{0, -cx}. Preprocessing of images
[0198] To normalize pixel values based on regional statistics, the fluorescence images were preprocessed, but the QBAM images were not. This normalization process 1) allows for the scaling of images for faster training of the DNN as the training time is reduced by the scaling of image values, 2) helps to remove some of the local background fluorescence or improve the contrast in areas with insufficient staining, and 3) enables images with different bit depths or contrasts to be processed by the same network. Since the pixel values are on an absolute scale and usually range from 0 to 1, the QBAM images were not normalized before feeding them into the network. Using the integral image, the local mean and standard deviation within a 127×127 pixel box centered on each pixel in the image were calculated. Each pixel was normalized according to Equation 13 based on the local pixel mean and standard deviation:
Equation
[0199] Where p i,j is the pixel value at row i and column j, μ i,j and σ i,j are the mean and standard deviation of the 127×127 pixel region centered on pixel p i,j and z i,j is the normalized pixel value. Assignment of pixel weights for DNN training
[0200] To improve the accuracy of the DNN, the training weight of each pixel in an image was adjusted based on the normalized pixel value according to the classification of each pixel. This adjustment helped the DNN to correctly label the cell boundaries in regions with poor signals. The pixels of the cell boundary were μi,j and σ i,j are the mean and standard deviation of only the boundary pixels in the 127×127 pixel region around p, and were normalized as described in Equation 1, except that. Then, for all boundary pixels where z i,j, >-1, a weight of 1 was assigned, and for all other pixels, a weight of -z i,j was assigned. This resulted in all bright pixels having a weight of 1, and for dark pixels, a weight inversely proportional to their pixel intensity was assigned. The weight for the cell body was assigned based on proximity to the cell boundary. The distance transform for each cell body was performed with respect to the cell boundary, and then the square root of that distance was used as the weight for pixel training. The weights of the cell bodies were truncated so that a value of 10 was assigned to all weights exceeding 10. i,j
[0201] The training weights were applied to the images during training using two different methods. The first method was to apply the weights as usual, multiplying the weights by the loss during error backpropagation. The second method was to use the weights as the pixel class label c in the logistic regression function (Equation 1). In both cases of the first method and the second method, c>0 for the cell boundary and c<0 for the cell body. In the first method, c = ±1, and in the second method, c was equal to the weight of the pixel. The second method was found to be better for several reasons. First, the second method resulted in a higher F1 score in the test dataset. Second, no deviation in the F1 score (which could suggest overfitting) for the test dataset and the training dataset was observed while training the DNN. Segmentation of QBAM Images Determination of Cell Boundaries in QBAM Images for DNN Training
[0202] A deep neural network (DNN-S) for segmenting QBAM images was developed. A subset of the samples was imaged using QBAM imaging in addition to fluorescence imaging of ZO-1, specifically in patients with AMD. Fluorescence images were captured together with transmitted light images of the iRPE. Next, images of the same region were captured using QBAM imaging. The QBAM images were registered to the fluorescence images of the iRPE by finding 256×256 pixel regions that had at least 97% correlation between the transmitted light images captured during fluorescence imaging and QBAM imaging. The correlation was evaluated using generalized normalized cross-correlation using the fast Fourier transform. The fluorescence images were segmented using the DNN, and that DNN segmentation was used as the cell boundary label for the QBAM images. DNN ensemble segmentation of QBAM images
[0203] Unlike the DNN used to segment fluorescence images of the RPE that achieved an F1 score of over 80%, the DNN for segmenting QBAM images achieved an F1 score of ≈60%. This was partly attributed to the fact that the ZO-1 boundary fluorescence did not overlap much with the boundaries observed in the QBAM images of some regions of the RPE monolayer. To improve the segmentation accuracy of DNN-S, an ensemble approach was used, training seven different networks and using the consensus of all the networks to determine the RPE boundary in the QBAM images, achieving an F1 score of 66%.
[0204] No preprocessing of the QBAM images was performed before processing them with DNN-S. The structure of DNN-S was the same as the structure used to segment fluorescence images, except that the input image had 3 channels instead of 1 channel. Those 3 channels were 3 captured QBAM images of the same field of view, but each QBAM image was captured at a different wavelength (488 nm, 561 nm, and 633 nm). Preprocessing of segmentation for feature extraction File Selection for Healthy-2 and AMD
[0205] From the directory structure containing all red, green, and blue brightfield images, calibration images, and live images, absorbance images derived from the blue channel, and ZO-1 fluorescence label images. For the Healthy-2 dataset (81,646 files in 1,564 folders), 2,580 absorbance files with a size of 1388 pixels × 1040 pixels were selected for segmentation. For the AMD dataset, 20 absorbance images with a size of 2,700 pixels × 3,000 pixels were selected. Color Inversion and Thinning (Skeletonize) of Segmentation Results
[0206] To remove artifacts and fill small gaps, a morphological binary operation of filling was applied to the segmented image containing the cell boundary lines. Thereafter, the segmentation was thinned so that the cell boundary lines became 1 pixel thick. Next, those images were inverted so that the foreground corresponded to the inside of the cells and the background corresponded to the cell boundaries. Finally, the binary inverted images were labeled so that each cell region was labeled with a unique id. Removal of Peripheral Cells in the Image
[0207] To calculate accurate cell-level features from the absorbance images on the reference segmentation mask, each mask was preprocessed to remove segments that reached the image boundary. This preprocessing step eliminates feature values that may correspond to partial cells. When evaluating the computer segmentation image for segmentation accuracy, the same preprocessing routine was applied to that computer segmentation image. Image Segmentation
[0208] In the case of the AMD dataset, the FOV stitched to a large image poses a scalability challenge for RAM during DNN-based model training. The size of the FOV varies depending on the experimental collection. Therefore, the collected images were pre-processed by stitching them into a large mosaic image and then dividing that mosaic image into 512×512 image tiles with 0% overlap. Feature extraction of images Feature extraction of vitiligo iRPE
[0209] From a directory structure containing 1,978 files in 439 folders, including all red, green, and blue brightfield images, calibration images and live images, absorbance images derived from the blue channel and ZO-1 fluorescence label images, 381 absorbance files derived from the blue channel were selected for image-level feature extraction. These files correspond to 3×3 fields of view (FOV) with 10% overlap. Due to the fact that some absorbance values were less than zero (calibration artifacts), a binary mask was created for each image to exclude those negative values from feature extraction. Using the Web Image Processing Pipeline (WIPP), a feature vector was calculated for each absorbance image (FOV) on the corresponding mask. The feature vectors originally included 14 intensity-based image features and 5 texture-based image features implemented in MATLAB® (R2017a). The intensity-based features correspond to the central moment ratio and entropy characteristics. The texture-based features are derived from the gray-level co-occurrence matrix (GLCM), which includes contrast, homogeneity, correlation, energy, and entropy. Scatter plots of the features were created using scripts in the R and ggplot packages, and the selected plots were included in the main manuscript. Feature extraction of AMD and healthy iRPE
[0210] To extract as many cell-level features as possible, multiple widely used image libraries for feature extraction integrated into WIPP were utilized. Feature values for the same feature definition can vary, but they provide insights into the amount of variation introduced not only by image acquisition but also by image processing. Intensity, shape, and texture features were calculated using the feature extractors in WIPP (Bajcsy et al., Web Microanalysis of Big Image Data, Springer International Publishing (2018b)), specifically, MATLAB® (2017a) and CellProfiler (Kamentsky et al., Improved structure, function and compatibility for CellProfiler: modular high-throughput image analysis software, Bioinformatics 27, 1179 - 11802011(2011)) applied to each cell region. Features corresponding to cell regions located at the image boundary and very large connected regions were discarded. Table 4 below summarizes these features: Table 4 List of cell feature cluster types extracted for each cell region and the software that derived those features
Table 4-1
Table 4-2
[0211] Table 4 shows 40 cell features, but since many features have subsets such as the angle or number of moments, the total number of measured cell features exceeded 315. Machine learning models for assay prediction
[0212] A machine learning-based regression model was constructed for cell-level features extracted from RPE cells, and the trained model could be used to predict TER and VEGF-Ratio measurements. All of those models were tested and trained on the same dataset for comparability. For the prediction of TER for healthy donor strains treated with afidicolin or HPI4, the data was split into test and training data by assigning 5 culture plates as training data and the 6th as test data (2 replicates of each treatment per plate). A similar procedure was done for VEGF-Ratio prediction, except that a subset of wells from plate 5 was selected and excluded from the training set and used as the test set (1 well from each treatment). The reason for this was that ELISA was not performed on the supernatants of all wells, and thus only a subset of wells was available for this analysis. To eliminate bias in sample selection, the selection of these plates / wells for training / testing was done blindly before data acquisition or analysis.
[0213] For AMD cell lines, SML models were trained from scratch for different combinations of iRPE monolayers. A total of 18 unique training image subsets were formed, where each image subset included test data with images of 1 clone from each donor. Using the average performance across all 18 subsets, important features were identified regardless of the machine learning method used or the combination of clones used in the test. Table 5 shows in detail all combinations not included in the training data. Table 5 List of 18 different clone subsets tested
Table 5-1
Table 5-2
[0214] The characteristics of RPE cells were extracted on a cell-by-cell basis. These characteristics include intensity, shape, and texture properties. Once extracted, the mean value, standard deviation, skewness, and kurtosis of each feature were calculated for each cell set within each image segmentation (described in the "Image Segmentation" section above). A subset of features was found to be highly correlated and have various amplitude ranges. Feature preprocessing was performed to extract the highly correlated features and normalize their dynamic range. This preprocessing was applied to all features and consisted of (1) z-normalization and (2) correlation-based feature selection (CFS) method that extracts all features with a correlation of more than 99.5%. This was done independently for all datasets within this manuscript. Thus, the highly correlated features within the combined dataset of AMD-iRPE strain and Healthy-2 strain were analyzed for co-correlation separately from the dataset used for AMD-iRPE sample identity (analyzed for co-correlation separately from the dataset of only Healthy-2). As a result, the total number of features used in each of these models was different. The normalization and extraction of highly correlated variables were also performed on the training / test data independently from other data. To evaluate the ability of DNN-S to accurately segment individual cell features and thus its most granular representation of the data, for a particular data, z-score regularization and feature correlation were performed on individual cell metrics rather than on the mean value, standard deviation, skewness, and kurtosis of the cells within the region of interest. Thus, the number of metrics for this analysis was considerably less than that for other models.
[0215] Among many ML models for classification and prediction, a subset of ML models that predict continuous variables such as TER or VEGF-Ratio measurements was considered. This subset of ML models includes multi-layer perceptron (MLP), linear support vector machine (L-SVM), random forest (RF), partial least squares regression (PLSR), and ridge regression (RR). For the donor classification task, the same models were chosen (except for ridge regression which does not have a classification form). Thus, for classification, ridge regression was replaced with a naive Bayes (NB) classifier. For all models, the feature weights were scaled based on their absolute magnitudes from 0 to 1, and then the average of all methods was obtained to determine which features have the highest relative weights compared to the prediction. Features with high averages suggest that these important features are consistently identified (i.e., model-independent of ML) and are relatively important for prediction. All models were trained using 10 - 30k-fold cross-validation and then tested on "dropout" data that was not in the training set or validation set within the k-fold cross-validation. Multi-layer perceptron (MLP)
[0216] The MLP model approximates the non-linear relationship between independent and dependent variables. In the multi-layer perceptron model, cell-level features were regarded as independent variables and the TER / VEGF values were regarded as dependent variables. An MLP model with 4 layers (2 hidden layers) was used, with 70 neurons used in the first hidden layer of the model and 40 neurons used in the second hidden layer. The number of hidden layers and the number of neurons in each hidden layer were empirically chosen so that the model does not overfit the input data and is not too simple to robustly fit the input data. To identify the features important for TER / VEGF prediction, Garson's algorithm was used (Garson, Interpreting neural-network connection weights, AI Expert 6, 46+(1991)) to calculate the weights of the relationship. Linear support vector machine (L-SVM)
[0217] The support vector machine constructs a hyperplane in a high-dimensional space to maximize the separation between data. In the case of a linear support vector machine (L-SVM), the plane is linearly correlated with each dimension / feature in that high-dimensional space, and the weights of the L-SVM can be directly converted into weights / importance for the features. The SVM function has terms that optimize the penalization of the hyperplane with respect to both the "closeness" (cost) of the fit to the data and the distance (epsilon) at which penalization occurs. Therefore, during model training, the cost and epsilon were iterated over ranges of 0.001 to 1.2 and 0.001 to 4096, respectively. All model optimizations were performed in R using the Liblinear package. Random Forest (RF)
[0218] Random Forest uses a "forest" or ensemble of decision trees to predict a factor. A decision tree can be thought of as a flow diagram for decision-making. Decision trees are constructed from an array of class-labeled training, where each node represents a test on an attribute, each branch represents the result of a test, and each terminal node classifies or performs a regression estimate. Random Forest combines multiple deep decision trees trained on different parts of the same training set with the goal of reducing the variance in the output of the decision trees. Therefore, the number of trees and the depth at which each tree branches can be optimized. Here, 125 trees were evaluated over 5 to 1500 branches to measure the optimal model performance. All model optimizations were performed in R using the Caret package. The relative importance of the variables was calculated as shown in Liaw, A. et al., Classification and Regression by randomForest, R News 2, 18 - 22 (2002). Partial Least Squares Regression (PLSR)
[0219] PLSR creates a linear regression model by projecting the predicted variables (TER and VEGF-Ratio) and observable variables (features of cell images) into a low-dimensional space. The PLSR model searches for a multidimensional direction within the "X" space (predicted variables) that explains the maximum multidimensional dispersion direction in the Y space (observable variables). In this way, by using high-dimensional data, real-world results can be predicted. The only variable optimized for PLSR is the number of dimensional components required to predict the desired result. For the model optimized in this report, the components were varied from 1 to 50 to determine which had the highest predictive power. The optimization of the model was performed in R using the Caret and PLS packages. The importance of a variable was defined as the sum of the absolute values of the feature coefficients within each dimensional component, multiplied by the percentage of the total variance that each component explained in the raw data. Ridge Regression (RR)
[0220] RR is a special form of ordinary least squares regression, where its prediction model is optimized to minimize the sum of the squared residuals between the predicted and actual results, and includes a gamma function for removing collinear, redundant, or confounding variables through optimization using L2 regularization. Similar to ordinary least squares regression, the importance of a variable can be determined from the absolute value of the weights of the feature coefficients. The model was optimized to reduce the number of variables to less than a total of 25 features. The optimization of the model was performed in R using the Caret and foba packages. DNN model for classifying donor identity
[0221] Two QBAM images of iRPE were taken, and a DNN was created to determine whether the iRPEs were from the same donor. The input was two images of 1024×1024×3 pixels. The basic layer structure consisted of a low-rank expansion as reported by Jaderberg et al., Speeding up Convolutional Neural Networks with Low Rank Expansions, ArXiv14053866 Cs(2014). The low-rank expansion layer consisted of a 1×3 convolutional layer followed by a 3×1 convolutional layer. The first layer had 8 neurons, and each subsequent layer doubled the number of neurons compared to the previous layer. After each convolutional operation, a leaky ReLU layer with a leak value of 0.1 followed. Each low-rank expansion layer was batch-normalized, and then followed by a 3×3 max-pooling layer with a stride of 2. Except for the first layer, a residual layer was added before each max-pooling operation, where the residual layer consisted of a 1×1 convolutional operation to scale the input layer to the same size as the output layer. The last layer was followed by a 3×3 average-pooling layer, and then a fully-connected layer or a dense convolutional layer (size = 64×64×256). The best network was judged based on the F1 score. Thanks to the simplicity and small size of this network, the network could be trained on the same 18 training / test datasets described in the following section. Machine learning model for classifying donor identity and outliers
[0222] A machine learning-based classification model was constructed for cell-level features extracted from RPE cells, and the trained model could be used to predict clonal outliers or donor identity. All models except the ridge regression model were evaluated for their ability to correctly classify donor identity. Instead of ridge regression, a naive Bayes model was performed. In the case of predicting clonal outliers, all data were subjected to PCA and clustering of iRPE strains using a hierarchical clustering method.
[0223] For AMD cell lines, the SML model was trained from scratch for different combinations of iRPE monolayers. A total of 18 unique training image subsets were formed, and each image subset included test data with images of one clone from each donor. Table 5 details all combinations not included in the training data. This approach used images of donor-derived clones to attempt to predict the "parent" donor strain of clones the network had never seen before. Principal component analysis (PCA)
[0224] PCA and hierarchical clustering were used to identify developmental outliers. PCA uses an orthogonal transformation to find the hyperplane of maximum variance that best describes the variables by transforming a set of correlated variables into a set of linearly uncorrelated variables called principal components, a reduced set of dimensionality values. In a large dimensional space (e.g., a large number of features), PCA is useful for reducing the dimensionality of the data to determine whether the features of cell images can generally classify different donors / clones relative to each other. Images were grouped at the clone level and day level, and for each clone / day combination, the mean value, standard deviation, skewness, and kurtosis of each feature were calculated. The features of these aggregated values were subjected to PCA. Principal components 1 and 2 were considered preferred indicators of the variability across the samples as they accounted for over 75% of the total variance in the data. The importance of a feature was defined as the sum of the absolute values of each feature coefficient within each dimensional component multiplied by the percentage of the total variance in the raw data explained by each component. PCA was performed using base R. Hierarchical clustering
[0225] Hierarchical clustering was performed on the PCA output for which the mean value, standard deviation, skewness, and kurtosis of each individual feature were calculated for each clone. In the case of hierarchical clustering, the Euclidean distance between all clones was calculated across all principal components, and the complete linkage distance was used for clustering. The splitting heights were selected in the three branches representing the three donors used in this study. Clustering was performed using base R functions. Naive Bayes
[0226] The Naive Bayes model describes the probability of a cell type being derived from a donor based on the prior probability of a cell derived from a donor having a given set of features. The Naive Bayes model assumes strong (naive) independence between features. That is, the Naive Bayes classifier considers each of the features of the cell image to contribute independently to the probability that the cell is derived from a given donor, regardless of the possible correlations between the features. The model was optimized for the required amount of Laplacian correction from 0.0 to 0.1, and also for whether a normal density or a kernel density was required for those features. The feature importance was obtained from the absolute value of the feature coefficients in the best-fit model. The model was optimized using the Caret and klaR packages in R. Quantitative and statistical analysis
[0227] All significance between groups shown for the vitiligo strain was performed using a linear mixed-effects model that controls for repeated measurements over time from a single well and multiple images taken for each well. These models were evaluated using the multicomp and nlme packages in R. R 2 Values, confidence intervals, and Kolmogorov–Smirnov statistics were calculated in base R.
[0228] In certain illustrative embodiments described above, it should be recognized that the various non-limiting embodiments described herein may be used separately, combined, or selectively combined for a particular application. Further, some of the various features of the non-limiting embodiments described above may be used without the corresponding use of other described features. Accordingly, the foregoing description should be regarded as merely illustrative of the principles, teachings, and exemplary embodiments of the present disclosure, and not as limiting thereof.
[0229] The arrangements described above should be understood to be merely illustrative of the application of the principles of the illustrated embodiments. Numerous modifications and alternative arrangements may be devised by those skilled in the art without departing from the scope of the illustrated embodiments, and the appended claims are intended to encompass such modifications and arrangements. The present invention provides, for example, the following items. (Item 1) A method for non-invasively predicting the characteristics of one or more cells and cell derivatives, the method comprising: training a machine learning model using at least one of a plurality of training cell images representing a plurality of cells and data identifying the characteristics of the plurality of cells; receiving at least one test cell image representing at least one test cell to be evaluated, the at least one test cell image being non-invasively acquired and based on absorbance as an absolute scale of light; providing the at least one test cell image to the trained machine learning model; predicting the characteristics of the at least one test cell using machine learning based on the trained machine learning model; and creating release criteria for a clinical cell preparation based on the predicted characteristics of the at least one test cell by the trained machine learning model comprising a method. (Item 2) The machine learning is performed using a deep neural network, and the method further includes a step of segmenting an image of the at least one test image into individual cells by the deep neural network, the method according to item 1. (Item 3) The machine learning is performed using a deep neural network, and the method further includes a step of classifying the at least one test cell based on the characteristics, the method according to item 1. (Item 4) The step of predicting further includes determining at least one of cell identity, cell function, effect of the delivered drug, disease state, and similarity to a technical replicate or previously used sample based on the step of classifying, the method according to item 3. (Item 5) The step of predicting the characteristics of the at least one test cell can be performed on either a single cell or a plurality of cells in one field of view in the at least one test cell image, the method according to item 1. (Item 6) The method further includes a step of visually extracting at least one feature from the at least one test cell image, wherein the step of training the machine learning model is performed using the at least one extracted feature, and the step of predicting includes identifying at least one feature of the at least one test cell using the trained machine learning model trained using the at least one extracted feature, and predicting the characteristics of the at least one test cell using the identified at least one feature, the method according to item 1. (Item 7) The at least one test cell image is acquired using quantitative brightfield absorbance microscopy (QBAM), the method according to item 1. (Item 8) The method includes receiving at least one microscopic image captured by a microscope; and converting the pixel intensity of the at least one microscopic image into an absorbance value The method according to item 7, further comprising (Item 9) wherein the method comprises a step of calculating the absorbance reliability of the absorbance value; a step of establishing the balance of the microscope by benchmarking; and a step of filtering colors when acquiring the image The method according to item 7, further comprising at least one of the above. (Item 10) The method according to item 3, wherein the first image processing of the at least one test image is performed by the deep neural network. (Item 11) The machine learning is performed using a deep neural network, and the method further comprises a step of segmenting the image of the at least one test image into individual cells by the deep neural network, and the feature is visually extracted from the segmented individual cells, as described in the item. (Item 12) The step of predicting the characteristics of the at least one test cell includes at least one of predicting the polarized secretion of transepithelial resistance (TER) and / or vascular endothelial growth factor (VEGF), predicting the function of the at least one test cell, predicting the maturity of the at least one test cell, predicting whether the at least one test cell is derived from a known donor, and determining whether the at least one test cell is an outlier with respect to a known classification, as described in item 1. (Item 13) The step of creating the release criteria for the clinical cell preparation further includes creating release criteria for drug discovery or drug toxicity by the trained machine learning model, as described in item 1. (Item 14) The method according to item 1, wherein the plurality of cells and the at least one test cell include at least one of embryonic stem cells (ESC), induced pluripotent stem cells (iPSC), neural stem cells (NSC), retinal pigment epithelial stem cells (RPESC), mesenchymal stem cells (MSC), hematopoietic stem cells (HSC), and cancer stem cells (CSC). (Item 15) The method according to item 15, wherein the first plurality of cells and the at least one test cell are derived from at least one of a plurality of ESC, iPSC, NSC, RPESC, MSC, or HSC or any cells derived therefrom. (Item 16) The method according to item 15, wherein the first plurality of cells and the at least one test cell include the major cell types derived from human tissue or animal tissue. (Item 17) The method according to item 1, wherein the identified and predicted characteristics include at least one of physiological characteristics, molecular characteristics, cellular characteristics, and / or biochemical characteristics. (Item 18) The method according to item 6, wherein the at least one extracted feature includes at least one of a cell boundary, a cell shape, and a plurality of texture metrics. (Item 19) The method according to item 18, wherein the plurality of texture metrics includes features within a plurality of cells. (Item 20) The method according to item 1, wherein the plurality of cell images and the at least one test cell image include a fluorescence image, a chemiluminescence image, a radiographic image, or a bright-field image. (Item 21) A step of determining whether the test image of the at least one test image is a large image; A step of dividing the large image into at least two tiles; A step of individually providing each of the tiles to the trained machine model; A step of combining the outputs of the processing by the trained machine model related to each of the tiles; and Providing the combined output for predicting the characteristics of the at least one test cell and making it a single output corresponding to the large image The method according to item 1, further comprising. (Item 22) A computing system for non-invasively predicting the characteristics of one or more cells and cell derivatives, wherein the computing system comprises: A memory configured to store instructions; A processor arranged to communicate with the memory Comprising, wherein the processor, when executing the instructions, Is configured to train a machine learning model using at least one of a plurality of training cell images representing a plurality of cells and data identifying characteristics for the plurality of cells; Is configured to receive at least one test cell image representing at least one test cell to be evaluated, the at least one test cell image being acquired non-invasively and based on absorbance as an absolute scale of light; Is configured to provide the at least one test cell image to the trained machine learning model; Is configured to predict the characteristics of the at least one test cell using machine learning based on the trained machine learning model; Based on the predicted characteristics of the at least one test cell, a release criterion for a clinical cell preparation is configured to be created by the trained machine learning model. Computing system.
Claims
[Claim 1] The invention described in the specification.
Citation Information
Patent Citations
RAPID AUTOMATED SCREENING SYSTEM AND METHOD FOR CELLS
JP2005525550A
Method for constructing prediction model for predicting quality of cell, program for constructing prediction model, recording medium recording the program, and device for constructing prediction model
JP2009044974A
Cell evaluation device, incubator, program, and culture method
JP2011229410A