Toxicity prediction of compounds in cellular structures
By using ML models to predict phenotypic characteristics and embeddings of cell structures, comparing these embeddings with known toxic compounds, the problem of difficult to reliably identify the toxicity of compounds in the prior art is solved, and more efficient and accurate toxicity prediction is achieved.
Patent Information
- Application Number
- CN202380073060.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-17
- Filing Date
- 2023-10-13
- Publication Date
- 2025-05-27
AI Technical Summary
Existing semi-automatic toxicity prediction assays are difficult to reliably identify the toxicity of compounds, especially on the cellular structure of the organ, resulting in an increased risk of performing in vivo tests on compounds that have passed in vitro tests.
Using a computer-implemented approach, by receiving an image set associated with multiple samples, the cellular structural phenotypic features within the sample is predicted using a first ML model, and then inputting these features into the second ML model to generate lower-dimensional phenotypic feature embeddings, comparing the distance between these embeddings and embeddings of known toxic compounds to predict the toxicity of the compound.
Improves reliable prediction of compound toxicity, reduces the risk of performing in vivo tests on compounds, and enhances the efficiency and accuracy of toxicity prediction.
Smart Images

Figure CN120051812A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to devices, systems, and methods for predicting the toxicity of compounds in the cellular structures of microscopy assay samples. Background Art
[0002] Cellular structures have been developed that can mimic and / or simulate the processes and / or functions of organs of a subject or patient. These cellular structures can be used for in vitro testing of the efficacy and / or toxicity of various compounds related to organs, rather than in vivo testing. Such cellular structures include immortalized cell lines that have been developed to mimic or simulate a particular organ of a subject. This has enabled semi - automated toxicity prediction test systems and methods that can be used to identify compounds that are toxic to the organs of a subject by using image microscopy and observing changes due to toxicity through dose - response (DR) charts and the like.
[0003] Conventional semi - automated test systems and methods can be used to identify compounds that produce a detectable signal to evaluate the effect of the compound on the cellular structure related to an organ. These test systems are referred to as assays. Once an assay for toxicity prediction has been developed, researchers can use the assay to identify compounds that have the desired activity related to toxicity. Typically, compounds will be tested at multiple concentrations, imaged using image microscopy, and a DR chart or other metric useful for the researcher to determine its toxicity can be generated. For example, analysis of the DR chart can allow the researcher to determine whether the compound is active and / or toxic, and at what concentration.
[0004] It is desirable to test a large number of potential compounds, which is typically done using high - throughput screening (HTS). This uses robots, data processing / control and imaging software, liquid - handling devices, and sensitive detectors, and allows researchers to quickly perform thousands or even millions of screening tests. However, the large amount of data generated at the imaging and DR steps of HTS activities requires careful analysis by the researcher to detect artifacts and correct erroneous data points before validation experiments.
[0005] Unfortunately, it has been found that even semi - automated toxicity prediction assays using HTS cannot reliably identify the toxicity of every compound (even compounds known to be toxic when analyzed on cell structures that mimic / organs). This increases the risk of performing in - vivo tests on compounds that have passed such semi - automated toxicity prediction assays. For example, 20% - 40% of patients with drug - induced liver injury (DILI) present a cholestatic and / or mixed hepatocyte / cholestatic injury pattern. Drug - induced hepatotoxicity or DILI is an acute or chronic response to natural or manufactured compounds. Conventionally, DILI can be classified based on clinical manifestations (hepatocyte, cholestatic or mixed), hepatotoxicity mechanisms, or histological appearance from a liver biopsy. Thus, reliable in - vitro toxicity prediction of compounds is an important part of drug / compound discovery or research programs.
[0006] There is a desire for improved methods, devices, systems, and / or architectures that can efficiently and reliably detect or predict the toxicity of compounds to cell structures, such as using in - vitro HTS assays. Summary of the Invention
[0007] According to a first aspect, there is provided a computer - implemented method for predicting the toxicity of one or more compounds applied to a plurality of samples of a cell structure in an in - vitro microscopy assay, the method comprising: receiving a set of images associated with the plurality of samples; inputting each image in the set of images into a first ML model, the first ML model being configured to predict phenotypic characteristics of the cell structure within the sample associated with each image; inputting each predicted phenotypic characteristic among the predicted phenotypic characteristics associated with each sample into a second ML model, the second ML model being configured to predict a lower - dimensional phenotypic characteristic embedding of each sample; comparing the distance between the lower - dimensional phenotypic characteristic embedding of each sample and the lower - dimensional phenotypic characteristic embedding of a sample to which a compound with a known toxicity has been applied; and based on the comparison, outputting an indication of the toxicity of each sample and the compound applied to the sample for each sample.
[0008] The computer - implemented method according to the first aspect, wherein the first ML model is a neural network or a convolutional neural network (CNN) model. As an option, the neural network or CNN model is trained using cell image training data for classification. As another option, the predicted phenotypic characteristics are embedded within a complete layer of the trained neural network or CNN model, and the method includes outputting the phenotypic characteristics from the complete layer. Optionally, the final complete layer of the neural network or CNN model is used to output the embedding of the phenotypic characteristics.
[0009] The computer-implemented method as described in the first aspect, wherein the second ML model is based on the Uniform Manifold Approximation and Projection (UMAP) algorithm or the t-SNE algorithm for dimensionality reduction of the phenotypic feature embeddings for the samples, and wherein the phenotypic feature embeddings are mapped to a lower-dimensional vector space for comparing the distances between the phenotypic feature embeddings of the samples with known toxicity compounds.
[0010] As an option, the second ML model is trained using unsupervised training based on the negative control samples and positive control samples among the multiple samples on the UMAP technique for predicting the toxicity distance metric associated with the samples to which compounds with known toxicity are applied. As another option, training the second ML model includes iteratively performing a grid search on a set of hyperparameters of the UMAP technique to select those hyperparameters that maximize the difference between the negative control samples and the positive control samples.
[0011] The computer-implemented method as described in the first aspect, wherein indicating the toxicity of the phenotypic feature embeddings of the samples to which compounds are applied includes applying the phenotypic feature embeddings of the samples to which compounds are applied to the second ML model to output the lower-dimensional embeddings of the samples to which compounds are applied; and determining an indication of the toxicity of the samples to which compounds are applied based on comparing the distances between the lower-dimensional embeddings and the embeddings of one or more samples to which compounds with known toxicity are applied.
[0012] The computer-implemented method as described in the first aspect, wherein indicating the toxicity of the phenotypic feature embeddings of the samples to which compounds are applied further includes: applying the phenotypic feature embeddings of the samples to which compounds are applied to the second ML model to output the lower-dimensional embeddings of the samples to which compounds are applied; and applying the lower-dimensional embeddings of the samples to which compounds are applied to a third ML model, which is trained to output an indication of the distance between the lower-dimensional embeddings and a set of lower-dimensional embeddings associated with the negative control samples.
[0013] The computer-implemented method as described in the first aspect further includes training the third ML model by performing a grid search on a set of hyperparameters of a high-dimensional distance metric algorithm, which maximizes the distance between the lower-dimensional embeddings of the negative control samples and the positive control samples while minimizing the distance between the lower-dimensional embeddings of the negative control samples or minimizing the distance between the lower-dimensional embeddings of the positive control samples.
[0014] The computer-implemented method as described in the first aspect, wherein the distance metric is the Wasserstein distance metric, and the high-dimensional distance metric algorithm is the Sinkhorn algorithm for estimating the Wasserstein distance between the embeddings.
[0015] The computer-implemented method as described in the first aspect, wherein the distance for comparing the toxicity of the phenotypic feature embeddings indicating the samples is based on the Wasserstein distance metric.
[0016] The computer-implemented method as described in the first aspect, wherein receiving the set of images associated with the plurality of samples further comprises: identifying viable samples of cell structures for analysis in an in vitro microscopy assay based on the following steps: automatically identifying a first set of samples available for analysis from the plurality of samples on an assay plate; generating a two-dimensional (2D) set of images for each sample in the first set of samples, the 2D set of images for each sample comprising a plurality of 2D image slices taken along the z-axis of each sample; and identifying a set of viable samples from the set of 2D image slices; and outputting data representing the set of viable samples for analysis as the set of images.
[0017] The computer-implemented method as described in the first aspect, wherein automatically identifying the first set of samples further comprises, for each sample in the plurality of samples: preprocessing the image of each sample; inputting the preprocessed sample image into a first machine learning (ML) model, the first ML model being configured to identify a region of interest (ROI) including cell structures in the input sample image; inputting the identified ROI of the sample image into a second ML model, the second ML model being configured to classify whether the sample is analyzable; and outputting the first set of samples including data representing those samples classified as analyzable.
[0018] The computer-implemented method as described in the first aspect, wherein the first ML model is a convolutional neural network (CNN) or other neural network trained to identify ROIs including cell structures, and the second ML model is a one-class support vector machine (SVM) configured to classify whether the ROI is analyzable.
[0019] The computer-implemented method as described in the first aspect, wherein the CNN is trained / configured based on a labeled training dataset, the labeled training dataset including a plurality of images, each of the images being annotated with a label including data representing whether there is a region of interest of cells and / or the position of the region of interest within the image, etc.
[0020] As an option, train / configure the one-class SVM for classifying whether the ROI is analyzable.
[0021] The computer-implemented method as described in the first aspect, wherein identifying the set of viable samples from the set of 2D image slices further comprises, for each sample: identifying the foreground, background, and multiple uncertain feature regions of the cell structure in each of the 2D image slices among these 2D image slices; iteratively combining the foreground, background, and the uncertain feature regions of these 2D image slices to generate a single 2D image of the cell structure; and selecting the sample for the set of viable samples based on the quality of the single 2D image.
[0022] As an option, the multiple uncertain feature regions include multiple uncertain foreground features and multiple uncertain background features.
[0023] The computer-implemented method as described in the first aspect, wherein the cell structure comprises one or more of the group consisting of: cell spheroid structures; vesicles; organoids; and any other suitable cell structure.
[0024] The computer-implemented method as described in the first aspect, wherein the plate includes a plurality of wells, and wherein each well has a sample of the cell structure.
[0025] According to the second aspect, there is provided an apparatus comprising a processor, a memory unit, and a communication interface, wherein the processor is connected to the memory unit and the communication interface, and wherein the processor and the memory are configured to implement the computer-implemented method according to the first aspect, its combinations, modifications thereof, and / or as described herein.
[0026] According to the third aspect, there is provided a computer-readable medium comprising data or instruction code that, when executed on a processor, causes the processor to implement the computer-implemented method according to the first aspect, its combinations, modifications thereof, and / or as described herein.
[0027] According to a fourth aspect, there is provided a tangible computer-readable medium comprising data or instruction code for predicting the toxicity of one or more compounds applied to a plurality of samples of cell structures in an in vitro microscopy assay, the data or instruction code causing at least one of one or more processors to perform at least one of the steps of the method when executed on the one or more processors: receiving a set of images associated with the plurality of samples; inputting each image in the set of images into a first ML model configured to predict phenotypic characteristics of cell structures within a sample associated with each image; inputting each predicted phenotypic characteristic among the predicted phenotypic characteristics associated with each sample into a second ML model configured to predict a lower-dimensional phenotypic characteristic embedding of each sample; comparing the distance between the lower-dimensional phenotypic characteristic embedding of each sample and the lower-dimensional phenotypic characteristic embedding of a sample of a compound with known toxicity; and based on the comparison, outputting an indication of the toxicity of each sample and the compound applied to the sample for each sample.
[0028] According to a fifth aspect, there is provided a system comprising: a receiver module configured to receive a set of images associated with the plurality of samples; a first ML model module configured to input each image in the set of images into a first ML model configured to predict phenotypic characteristics of cell structures within a sample associated with each image; a second ML model module configured to input each predicted phenotypic characteristic among the predicted phenotypic characteristics associated with each sample into a second ML model configured to predict a lower-dimensional phenotypic characteristic embedding of each sample; a distance comparison module configured to compare the distance between the lower-dimensional phenotypic characteristic embedding of each sample and the lower-dimensional phenotypic characteristic embedding of a sample of a compound with known toxicity; and an output module configured to output an indication of the toxicity of each sample and the compound applied to the sample for each sample based on the comparison.
[0029] In various embodiments, computer program instructions optionally stored on a non-transitory computer-readable medium, the computer program instructions causing a data processing device to execute the program instructions to cause the one or more processors to perform operations including one or more aspects of the embodiments described above and / or below (including one or more aspects of the appended claims).
[0030] In various embodiments, a device is disclosed that includes a computer-readable storage medium and one or more processors. The computer-readable storage medium has program instructions included therein, and the one or more processors are configured to execute the program instructions to cause the device to perform operations including one or more aspects of the embodiments described above and / or below (including one or more aspects of the appended claims). The device may include one or more processors or dedicated computing hardware. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] To make the present invention more readily understood, embodiments of the present invention will now be described by way of example only with reference to the accompanying drawings, in which:
[0032] FIG. 1a shows an example toxicity prediction system for predicting the toxicity of a compound for a cell structure of a sample applied to a microscopy assay according to some embodiments of the present invention;
[0033] FIG. 1b shows an example toxicity prediction process for the system of FIG. 1a according to some embodiments of the present invention;
[0034] Figure 2 Shows an example neural network for predicting a phenotypic embedding of a cell structure in a sample of the microscopy assay of FIG. 1a according to some embodiments of the present invention;
[0035] FIG. 3a shows an example assay plate according to some embodiments of the present invention, the assay plate having a group of negative and positive control samples for training a deep learning toxicity model and a group of test samples for input into the trained DL toxicity model;
[0036] FIG. 3b shows an example unsupervised training process of a deep learning toxicity model using the group of negative and positive control samples of FIG. 3a according to some embodiments of the present invention;
[0037] FIG. 3c shows an example of using the trained deep learning toxicity model of FIG. 3b to predict the toxicity of a compound according to some embodiments of the present invention;
[0038] FIG. 4a shows another example assay plate according to some embodiments of the present invention, the assay plate having a group of negative and positive control samples for training a deep learning toxicity model and a group of test samples for input into the trained DL toxicity model;
[0039] FIG. 4b shows an example distance matrix of negative and positive control samples of the trained DL model of FIG. 4a according to some embodiments of the present invention;
[0040] Figure 4c shows another example distance matrix of negative and positive control samples and test samples for predicting the toxicity of compounds of a test sample using the trained DL toxicity model of Figure 4b according to some embodiments of the present invention;
[0041] Figure 4d shows an example of a conventional toxicity prediction method applied to a set of 8 compounds for samples of cell structures;
[0042] Figure 4e shows an example of the conventional toxicity prediction results of a set of compounds of the conventional toxicity prediction method from Figure 4d;
[0043] Figure 4f shows an example of the toxicity prediction results of the trained DL toxicity model for the same set of compounds as in Figure 4e;
[0044] Figure 5a shows an example of a quality control image analysis process according to some embodiments of the present invention, the quality control image analysis process being used to select and preprocess microscopic images of cell structures before being used for training the DL toxicity model and / or as an input to the trained DL toxicity model, the trained DL toxicity model being used to predict the toxicity of compounds applied to the cell structures;
[0045] Figure 5b shows an example of a first quality control process of a first quality control device according to some embodiments of the present invention, the first quality control process being used together with the quality control image analysis process of Figure 5a to select a first set of viable images of a sample for analysis;
[0046] Figure 5c shows an example of an image analyzer of a first quality control device for use together with the first quality control process of Figure 5b according to some embodiments of the present invention;
[0047] Figure 5d shows an example of a first quality control device for use together with the first quality control process of Figure 5b according to some embodiments of the present invention;
[0048] Figure 5e shows an example of a second quality control process of a second quality control device according to some embodiments of the present invention, the second quality control process being used together with the quality control image analysis process of Figure 5a to preprocess and select a final set of viable images of a sample for analysis;
[0049] Figure 5f shows another example of a second quality control process controlled by a second quality control device according to some embodiments of the present invention, the second quality control process being used together with the quality control image analysis process of Figure 5a to preprocess and select a final set of viable images of a sample for analysis;
[0050] Figure 6a is a schematic diagram of a system / apparatus for performing the methods described herein; and
[0051] Figure 6b It is a schematic diagram of another example system for performing the methods described herein.
[0052] Throughout the drawings, common reference numerals are used to indicate like features. Detailed Description
[0053] The various example embodiments described herein relate to methods, devices, and systems for automatically, efficiently, and reliably testing and predicting the toxicity of compounds applied to samples of cellular structures in HTS microscopy assays. The toxicity prediction system receives a set of images of samples of cellular structures, where at least one set of the assayed samples has been perturbed by one or more test compounds. The received set of images is applied to a deep learning (DL) model that is trained and configured to predict the toxic effects of the compounds on the perturbed samples of cellular structures. The DL model generates high-dimensional phenotypic representations of the cellular structures represented in each of the received image samples. These high-dimensional phenotypic representations are further processed into lower-dimensional phenotypic embeddings, focusing on the features of the cellular structures associated with toxicity. A toxicity prediction for each image sample is output based on the distance between each lower-dimensional phenotypic embedding and a negative control lower-dimensional phenotypic embedding or a similarity metric used to estimate that distance, the negative control being associated with a sample of cellular structures that has not been perturbed by any of the test compounds.
[0054] The cellular structures of the samples are associated with an organ of a subject or patient and can be designed to mimic or simulate the organ. The cellular structures can be, but are not limited to, for example, based on at least one of the group consisting of: cell spheroids; vesicles; organoids; cellular structures of immortalized cell lines; and any other suitable cellular structures that mimic or simulate one or more processes of an organ of a subject or patient.
[0055] The toxicity prediction system is described herein with reference to drug-induced liver toxicity or drug-induced liver injury (DILI) only by way of example and not limitation as the acute or chronic response of the liver to natural or manufactured compounds. Up to 20%-40% of DILI patients present a cholestatic and / or mixed hepatocellular / cholestatic injury pattern. The assayed samples use cellular structures associated with the liver, which follow cholestasis into hepatocytes in vitro. The type of cellular structures used is, but is not limited to, for example, HepaRG (RTM) cells, which are terminally differentiated hepatocytes derived from a human hepatic progenitor cell line that retains many of the properties of primary human hepatocytes.
[0056] The HepaRG cell is an immortalized cell line with four main characteristics: 1 - the full array of functions, responses, and regulatory pathways of primary human hepatocytes, including: phase I and II, as well as transporter activities consistent with those found within the primary human hepatocyte population; 2 - the formation of bile canaliculi; 3 - the potential to express the main characteristics of stem cells; 4 - high plasticity and the ability to fully transdifferentiate. The cells can be directly used in experimental assay plates after 7 days of maturation. HepaRG cells can form cell spheroids that mimic or simulate one or more cellular processes of the liver, and these cell spheroids can be stained and / or fluoresced and imaged during in vitro microscopy assays for downstream analysis.
[0057] Although a toxicity prediction system is described herein with reference to hepatocyte structures / spheroids (e.g., the HepaRG cell line) and DILI, this is merely an example and the present invention is not limited thereto. Those skilled in the art should understand that the toxicity prediction system can be trained and applied to any type of cell structure, any cell line, cell array, and / or well of a cell sample to which a compound has been added, associated with any organ and / or any related disease. For example, the toxicity prediction system can be applied to cell structure samples that mimic organs, such as but not limited to, for example, lung, skin, kidney, pancreas, liver, heart cell structure / heart, nerve cell structure, and / or any other organ of a subject or patient. Such cell structures can be used in samples to which a compound has been applied in in vitro microscopy assays, etc., and the toxicity can be automatically analyzed by the toxicity prediction system.
[0058] Figure 1a shows an example toxicity prediction pipeline 100 for automatically predicting the toxicity of one or more compounds applied to a plurality of samples of a cell structure in an in vitro microscopy assay. The toxicity prediction pipeline 100 includes an in vitro HTS microscopy assay system 102, an imaging analysis system 104, and a toxicity prediction system 106. The in vitro HTS assay system 102 is configured to obtain a set of samples of a cell structure 102a and apply one or more compounds or reagents 102b to the set of samples of the cell structure 102a for input into a set of well samples in a microscopy assay plate 102c for HTS staining and microscopy assay imaging 102d. In the HTS staining and microscopy assay imaging 102d, the well samples can be treated / stained with a fluorescent reagent / compound to emphasize the cell structure of each sample. For example, the treated / stained cell structure (such as, for example, spheroid structures, vesicles, and / or cell nuclei) can be imaged by a microscopy imaging system. Thus, the in vitro HTS microscopy assay system 102 is configured to output a set of well image samples 103. Each well sample image in the set corresponds to the cell structure of the well sample in the assay and each sample to which a compound has been applied.
[0059] The imaging analysis system 104 can be configured to receive the set of well sample images and perform image preprocessing and / or analysis to identify which samples from the set of well samples are viable for further downstream analysis, such as, for example, toxicity prediction of compounds applied to the samples of the set of well samples. One or more image processing and / or machine learning algorithms can be applied to identify the viability of each well sample image based on any detected image artifacts and / or imaging defects, etc., and / or to enhance or emphasize the cell structures of interest within each viable well sample image in the set of well samples. Thus, the imaging analysis system 104 can output a set of viable well sample images 105 for further downstream analysis. Substantially, the set of viable well sample images 105 can be any suitable set of images of the cell structures from the set of well samples of the assay that sufficiently depicts the cell structures of the samples for automated analysis.
[0060] The toxicity prediction system 106 is configured to receive, during an in vitro microscopy assay, a set of images 105 of the cell structures from a plurality of cell structure samples to which one or more compounds have been applied. In this case, the received set of sample images 105 can be the set of viable well sample images 105 output from the imaging analysis system 104. However, the toxicity prediction system 106 can receive any suitable set of images that sufficiently depicts the cell structures of the set of well samples to which the compounds have been applied. For example, a set of well image samples 103 output from the in vitro HTS microscopy assay system 102 can be used, provided that the images are sufficiently artifact- and / or defect-free and the components of the cell structures of each sample are analyzable.
[0061] In the toxicity prediction system 106, the received set of sample images 105 can each be input into deep learning (DL) toxicity prediction models 106a - 106c, which are configured to predict the toxicity of each corresponding compound applied in assays from an in vitro HTS microscopy assay system 102. The DL toxicity prediction models 106a - 106c can be based on any one or more DL modeling techniques / algorithms and / or machine learning (ML) techniques / algorithms that have been used to train the DL toxicity models to identify or predict whether each received sample image in the received set of sample images 105 indicates toxicity, even when the compound has not been applied to the cell structure in one or more of these well samples. The one or more DL / ML techniques / algorithms can be based on supervised ML, unsupervised ML, and / or semi - supervised ML algorithms, among others. However, for the task of training the DL toxicity prediction models 106a - 106c, it has been found that supervised learning is difficult due to the limited number of labeled training data sets related to cell structures that indicate toxicity or not, depending on whether the compound has been applied. Thus, a combined supervised / unsupervised DL model training architecture can be used to train one or more component models of the DL toxicity prediction models 106a - 106c.
[0062] For example, in the example of FIG. 1a, the DL toxicity prediction models 106a - 106c can include a trained machine learning (ML) phenotypic feature extraction (FE) model 106a for extracting phenotypic features of cell structures from each received sample image in the received set of sample images 105. Supervised learning can be used to train the ML phenotypic FE model 106a. The ML phenotypic FE model 106a can be based on, but not limited to, for example, a neural network (NN) classifier that is trained using supervised training on an easily - obtainable labeled / annotated training data set to classify images of cells, organoids, spheroids, cell structures, etc. (e.g., classifying an image of a cell structure to determine whether the cell is a cancer cell or a tumor cell). The NN classifier can be based on any NN structure, such as but not limited to, for example, a feed - forward NN (FNN), a recurrent NN (RNN), an artificial NN (ANN), a convolutional NN (CNN), any other type of NN, modifications thereof, combinations thereof. Before the output classification layer or SoftMax output of the NN classifier, a phenotypic representation of the cell structure can be embedded through the high - dimensional output of one of the hidden layers or full layers of the NN classifier. The NN classifier is configured to output an embedding of the phenotypic representation of the cell structure from the hidden layer or full layer, rather than outputting a classification.
[0063] Typically, the phenotypic representation embedding of the input image samples from the received image sample set 105 is a high-dimensional representation of phenotypic features (e.g., for a CNN-type NN classifier / model, the dimension can be on the order of 2024 or greater). The trained NN classifier can be used to output a high-dimensional phenotypic feature representation 107 of the cell structure for each input well sample in the sample image set 105 to which a compound has been applied. As an example, the neural network classifier can be based on a convolutional neural network (CNN) and be used to classify images of cell structures (e.g., classify images of cell structures to determine whether the cells are cancer cells), where, once trained, one of the final complete layers of the CNN can be used as the phenotypic feature representation of each image sample in the received image sample set 105 input thereto.
[0064] The ML phenotypic embedding model 106a is coupled to a trained ML low-dimensional (LD) embedding model 106b that is trained and configured to optimally embed the high-dimensional phenotypic representation 107 of the predicted phenotypic features of each received sample image of the received image sample set 105 into a lower-dimensional phenotypic embedding 108 in a lower-dimensional space for further analysis, where the lower-dimensional phenotypic feature embedding 108 includes those phenotypic features associated with toxicity. Instead of using supervised training on the ML LD embedding model 106b, one or more unsupervised DL / ML techniques / algorithms are used to ensure that the trained ML LED embedding model 106b outputs an LD phenotypic embedding 108 representing the phenotypic features associated with the toxicity of the cell structure. The unsupervised DL / ML techniques / algorithms can be at least based on clustering algorithms or dimensionality reduction algorithms, such as but not limited to, for example, support vector machine (SVM), uniform manifold approximation and projection (UMAP), or t-distributed stochastic neighbor embedding (t-SNE)-type algorithms, combinations thereof, modifications thereto, etc.
[0065] For example, negative control samples and positive control samples included on the assay plate 102c can be used in an in vitro microscopy assay to train the ML LD embedding model 106b. That is, the assay plate 102c can include a negative control sample group and a positive control sample group, where the negative control sample group can be a first set of wells of the assay plate 102c where the samples are undisturbed or no compound has been applied thereto, and the positive control sample group can be a second set of wells of the assay plate 102c where the samples are perturbed with a known compound having known toxicity. The assay plate 102c can also include a third set of wells of the assay plate 102c that includes cell structure samples with a compound applied thereto. Thus, the negative image group and the control image group associated with the negative control samples and the positive control samples can be used to train the ML LD embedding model 106b in an unsupervised manner.
[0066] For example, the negative image group and the control image group are input into the ML phenotypic embedding model 106a, which outputs corresponding negative high-dimensional phenotypic representations and positive high-dimensional phenotypic representations. Then, iterative optimization of the parameters of the UMAP or t-SNE algorithm can be performed on the high-dimensional negative and positive comparison phenotypic representations output from the ML phenotypic embedding model 106a, where the difference between the obtained LD negative and positive comparison phenotypic representations embeddings is maximized. The obtained optimized parameters can be used together with the UMAP or t-SNE algorithm for dimensionality reduction of the high-dimensional phenotypic representation 107 corresponding to the third group of well samples. The LD phenotypic representation 107 can be output for comparison with the negative control (NC) LD phenotypic representation using a suitable distance or similarity metric used by the trained ML distance model 106c.
[0067] The ML LD embedding model 106b is coupled to the trained ML distance / prediction models 106c / 106d, which are configured to perform a distance or similarity metric comparison / estimation 109 between the LD phenotypic representation of one of the image samples of the well and the negative control LD phenotypic representation using a suitable distance or similarity metric related to the LD space of the LD phenotypic representation. The ML distance model 106c can output data representing each distance comparison estimate 109, which is input into the ML prediction unit 106d to predict toxicity. It should be noted that the LD space of the LD phenotypic representation embeddings output by the ML LD embedding model 106b can still be considered a high-dimensional space where the Euclidean distance metric or similarity metric decomposition and / or cannot be reliably used. Therefore, a high-dimensional distance metric or similarity metric can be used based on but not limited to, for example, the Wasserstein distance and / or any other high-dimensional distance metric or similarity metric. The ML distance / prediction models 106c / 106d output a toxicity prediction as a probability based on the distance comparison 109 (e.g., Wasserstein distance comparison) between each LD phenotypic representation of the sample and the negative control LD phenotypic representation.
[0068] Figure 1b shows an example toxicity prediction process 110 for a toxicity prediction system 106 to predict the toxicity of one or more compounds for multiple samples of cell structures in the in vitro microscopy assay pipeline 100 of Figure 1a. The toxicity prediction process 110 includes the following steps: In step 112, an image set associated with multiple samples is received based on the output of the in vitro microscopy assay. Each sample image includes image data that fully describes the cell structure of the associated sample for automated processing and analysis. In step 114, each image in the image set is input into a first ML model, which is configured to predict the phenotypic characteristics of the cell structure within the sample associated with each image. In step 116, each of the predicted phenotypic characteristics associated with each sample is input into a second ML model, which is configured to predict a lower-dimensional phenotypic characteristic embedding for each sample. In step 118, the distance between the lower-dimensional phenotypic characteristic embeddings of each sample and the lower-dimensional phenotypic characteristic embeddings of samples (e.g., negative control samples) to which a compound with known toxicity has been applied is compared. In step 120, based on the comparison, an indication (e.g., probability) of the toxicity of each sample and the compound applied to that sample is output for each sample.
[0069] Figure 2 Shows an example neural network classifier 200 in the ML phenotypic FE model 106a of the toxicity prediction system 106 for Figure 1a. The NN classifier 200 is configured to predict or output a phenotypic representation / embedding from an image of a cell structure in a sample of the microscopy assay of Figure 1a. The NN classifier 200 is based on a CNN-type architecture, which includes a first part of a CNN-type network 202, followed by one or more fully connected layers 204a - 204n and a classifier output layer 206. One or more of the outputs 205a - 205n of the fully connected layers 204a - 204n can be tapped and / or selected 208 for output as a phenotypic representation embedding 210 associated with the input sample image representing the cell structure. As an example, the NN classifier 200 can be a 50-layer RESNET CNN architecture with multiple fully connected layers (e.g., RESNET50(RTM) or VGG(RTM), etc.).
[0070] The NN classifier 200 is trained on image data containing cell structures relevant to a classification task, where relevant phenotypic information of the cell structures can be extracted from the output of one of the hidden layers or fully connected layers of the trained NN classifier 200. The NN classifier 200 can be pre-trained for the classification task using labeled image data associated with cell types / structures. For example, a labeled training image data set can be used to train the NN classifier 200 to classify images as but not limited to, for example, cancerous or non-cancerous, identification of cells / cell structures, and / or other diseases affecting cell function / structure.
[0071] Substantially, when trained for the classification task, the layers within the CNN architecture of the NN classifier 200 begin to "recognize" cell primitives, cell aggregates, how cells form tissue regions, etc. to identify larger cell macrostructures, which can be used to form a high-dimensional phenotypic representation of the cell structures present in the input image embedding the cell structure samples in the complete layer 204a - 204n. The classification task or output 206 is only used to train the NN classifier 200. As another example, the NN classifier 200 can be trained on labeled cell image data from ImageNet and then fine-tuned using a limited number of training image data items associated with cell data from microscopy assays, etc. and / or stained / fluorescent cell images.
[0072] The NN classifier 200 architecture can be configured to have multiple fully connected layers 204a - 204n before the output classification 206. One of the fully connected layers 204n is selected from all the fully connected layers 204a - 204n based on the layer determined to be focused on cell structures in terms of deep representation. As an option, the toxicity prediction system 106 can perform automatic optimization on each of the ML models 106a - 106c to determine which one of the fully connected layers 204a - 204n of the trained NN classifier 200 provides the best high-dimensional phenotypic representation output highlighting the toxic effects or toxicity in cell structures.
[0073] FIG. 3a is an exemplary assay plate 300 having a negative control sample well group 302 and a positive control sample well group 304, where samples therein are imaged during an in vitro microscopy assay for training a deep learning (DL) toxicity model 106. The assay plate 300 also includes a test sample well group 306, where samples therein are imaged during the in vitro microscopy assay for input into the trained DL toxicity model for predicting the toxicity of any compound applied to the test sample well group 306. The negative control sample well group 302 includes samples that have not been perturbed by the compound being tested or have been perturbed by a non-toxic compound. For example, in an in vitro microscopy assay using hepatic spheroids, dimethyl sulfoxide (DMSO) can be applied to the NC samples (e.g., 0.6% DMSO). The positive control sample well group 304 includes samples that have been perturbed by a compound having a known toxicity associated with the cellular structure used in the samples. For example, in an in vitro microscopy assay using hepatic spheroids, chlorpromazine (CPZ) at a concentration that produces a toxic effect (e.g., 400 μM) can be used to perturb the PC samples.
[0074] The DL toxicity model 106 is trained and configured to first extract phenotypic features of the cellular structure of each sample in the sample images and then estimate the phenotypic distance between the extracted phenotypic features and the phenotypic features associated with the samples in the negative control sample well group 302. Microscopy images of the negative control samples and positive control samples in the negative control sample well group 302 and the positive control sample well group 304 are used together with unsupervised DL / ML techniques to train the DL toxicity model 106 to estimate the toxicity of the compound of each sample using the phenotypic distance between the extracted phenotypic features and the phenotypic features associated with the samples in the negative control sample well group 302. For example, the negative control samples are used to establish the average and deep phenotypic representation embeddings of the negative control samples (e.g., the first 20 wells). Once the average and deep phenotypic representation embeddings of the negative control samples have been established, this can be used as a reference for determining the distance of the low-dimensional embeddings of subsequent test samples.
[0075] In this example, the assay plate 300 has a well array consisting of row and column wells (e.g., columns A - P and rows 01 - 24), which are mapped to a negative control sample well group 302, a positive control sample well group 304, and a test sample well group 306. In this example, the assay plate 300 has an array consisting of 16 columns and 24 rows of wells or a total of 384 wells. Although the assay plate 300 shows the negative control well group 302 as mapped to columns A - P and rows 1 - 8, the positive control well group 304 as mapped to columns A - P and rows 20 - 24, and the test sample well group 306 as mapped to columns A - P and rows 9 - 19, this is merely an example and the present invention is not limited thereto. Those skilled in the art should understand that any mapping can be defined to specify the mapping between the sample wells of the assay plate 300 and the negative control sample wells 302, positive control sample wells 304, and test sample well group 306. The mapping of samples to the wells in the assay plate 300 can be automatically defined by the software management system of the HTS in vitro microscopy assay system 102 according to the number of samples required in each group of negative control sample wells 302, positive control sample wells 304, and test sample wells 306. For example, the samples in the test sample well group can have one or more compounds at different concentrations applied to the sample. The HTS in vitro microscopy assay system 102 can be programmed to map the type and concentration of the compounds onto the assay plate 300, depending on the number of compounds to be tested, the number of different concentrations, and the number of replicates, etc.
[0076] Figure 3b shows an example unsupervised training process 310 for training a deep learning toxicity model, which includes the ML low - dimensional (LD) embedding model 106b and the ML distance model 106c of the toxicity prediction system 106 of Figure 1a. Training the DL toxicity model uses the negative control group sample wells 302 and positive control group sample wells 304 of the samples in Figure 3a. The unsupervised training process 310 includes the following steps:
[0077] In step 312, the training process 310 receives negative control (NC) and positive control (PC) phenotypic representation embeddings associated with the NC and PC samples corresponding to the NC sample well group 302 and PC sample well group 306 of the plate 300. The images based on the NC and PC samples corresponding to the NC sample well group 302 and PC sample well group 306 of the plate 300 can be input into the ML phenotypic feature extraction model 106a of Figure 1a, which ML phenotypic feature extraction model can be based on Figure 2 the NN classifier 200 and has been pre - trained as described with reference to Figures 1a - 2. In view of this, the ML phenotypic feature extraction model 106a outputs the corresponding NC and PC phenotypic representation embeddings 107. The NC and PC phenotypic representation embeddings 107 are high - dimensional embeddings (e.g., 2024 elements per image).
[0078] In step 313, joint unsupervised training is performed on the ML LD embedding model 106b and the ML distance / prediction models 106c / 106d based on the following steps: In step 314, unsupervised training is performed on the ML LD embedding model 106b to estimate the LD phenotypic embeddings 108 that represent the toxicity features associated with the NC and PC phenotypic representations embeddings. The unsupervised training can be performed by iteratively optimizing the ML LD embedding model 106b within the range of parameter sets associated with the ML techniques (e.g., UMAP or t-SNE) used to configure the ML LD embedding model 106b. The criterion for the ML technique used to adjust or generate the ML LD embedding model 106b is to find a subset of parameters within this range of parameter sets that maximizes the difference between the NC embeddings and the PC embeddings using the input high-dimensional NC and PC phenotypic representations embeddings. In each iteration, a subset of parameters is taken from this range of parameter sets and applied to the ML technique (e.g., UMAP or t-SNE), which adjusts or generates the ML LD embedding model 106b to output the LD NC and PC phenotypic embeddings.
[0079] As an example, the input high-dimensional NC and PC phenotypic representations embeddings 107 output from the ML phenotypic FE model 106a are high-dimensional vectors for each image (e.g., each embedding can be 2024 elements per image). The ML phenotypic FE model 106a has not been specifically trained for toxicity prediction, but rather for the biological structure or phenotypic representation of the cell structure within each image of the sample. Therefore, the resulting high-dimensional NC and PC phenotypic representations embeddings describe the phenotypic representation of each corresponding cell structure as much as possible within the dimension (e.g., 2024). The UMAP algorithm can be used to generate the ML LD embedding model 106b to not only reduce the dimension of the high-dimensional NC and PC phenotypic representations, but also ensure that the phenotypic representations associated with toxicity are retained or concentrated within the resulting LD NC and PC embedding vectors 108 (e.g., each embedding can be 64 elements per image) output from the ML LD embedding model 106b. This is performed by the UMAP algorithm finding the best set of parameters within this range of parameter sets that maximizes the difference between the high-dimensional NC phenotypic representation embeddings and the PC phenotypic representation embeddings. The ML LD embedding model 106b is trained to unbiasedly find the relevant biological structures associated with toxicity within each image without knowing the cell type or its application. The ML LE embedding model 106b can use the LD embeddings (e.g., 64 dimensions in the embedding) to represent the entire negative control group by taking the simple average profile of all negative controls. This can be used by the ML distance / prediction models 106c / 106d to determine the toxicity prediction based on estimating the distance or similarity between the LD embeddings of the test samples from the test sample well group 306 and the average LD NC embeddings of the NC samples from the NC sample well group 302 using a distance or similarity metric.
[0080] In step 315, the ML LD embedding model 106b outputs NC and PC LD phenotypic embeddings 108, where each embedding includes phenotypic information of the corresponding cellular structure related to toxicity. The NC and PC LD phenotypic embeddings are used as the input for step 316 to train the ML distance model 106c.
[0081] In step 316, unsupervised training is performed on the ML distance model 106c for estimating based on a suitable distance or similarity of high-dimensional vector metrics (e.g., Wasserstein distance) to maximize the distance between the NC LD phenotypic embedding and the PC LD phenotypic embedding. This means that the trained ML distance model 106c can determine the distance between the NC LD phenotypic embedding of the test sample 109 and the LD phenotypic embedding for determining the probability or indication of whether the compound associated with each test sample is toxic. The unsupervised training can be performed by iteratively optimizing the ML distance model 106c within the range of the parameter set associated with the ML distance algorithm / technique (e.g., Wasserstein distance algorithm, such as Earth Mover's distance and / or Sinkhorn distance algorithm) used to configure the ML distance model 106c. The criterion for adjusting or generating the ML distance algorithm / technique of the ML distance model 106c is to find a subset of the parameters of the ML distance algorithm that maximizes the distance between the set of NC LD embeddings and the set of PC embeddings, but minimizes the distance between the embeddings within the set of NC LD embeddings and minimizes the distance between the embeddings within the set of PC LD embeddings. In each iteration, a subset of parameters is selected from the range of the parameter set of the ML distance algorithm and applied to the ML distance algorithm (e.g., Earth Mover's distance algorithm or Sinkhorn algorithm), which adjusts or generates the ML distance model 106c / 106d for outputting a toxicity indication between the LD phenotypic embedding and the set of NC LD embeddings.
[0082] In step 317, it is determined whether the best distance between the set of NC LD embeddings and the set of PC LD embeddings has been obtained for the parameter set range of the ML distance model 106c. This can also include determining whether the difference between the NC LD embedding and the PC LD embedding has been maximized based on the parameter sets tested so far. If this is the case or if there are no more parameter combinations in the parameter sets of the available ML LD embedding model 106b or ML distance model 106c, the process 310 proceeds to step 319. Otherwise, the process proceeds to step 318 for further parameter selection and adjustment / training of the ML LD embedding model 106b and / or the ML distance model 106c.
[0083] In step 318, additional parameters are selected from the parameter set associated with the ML LD embedding model 106b for further training thereof, where process 310 proceeds to step 314. Similarly, additional parameters can be selected from the parameter set associated with the ML distance model 106c for further training thereof, where process 310 proceeds to step 316.
[0084] In step 319, the ML LD embedding model 106b selects a subset of parameters within the parameter set associated with the ML LD embedding model 106b that maximizes the difference between the NC high-dimensional phenotypic embedding and the PC high-dimensional phenotypic embedding for outputting an LD phenotypic embedding related to the high-dimensional phenotypic embedding of the test sample corresponding to the test sample well group 306. Similarly, the ML distance / prediction models 106c / 106d select a subset of parameters within the parameter set associated with the ML distance model 106c that maximizes the distance (e.g., Wasserstein distance) between the NC LD phenotypic embedding set and the PC LD phenotypic embedding set but also minimizes the distance within each NC and PC LD phenotypic embedding set for outputting a toxicity estimate based on the distance between the LD phenotypic embedding of the test sample corresponding to the test sample well group 306 and the NC phenotypic embedding set or the average NC phenotypic embedding.
[0085] FIG. 3c shows a toxicity prediction process 320 for predicting the toxicity of a compound using the trained deep learning toxicity model of FIG. 3b. From step 319 of the toxicity training process 310 of FIG. 3b, the ML LD embedding model 106b and the ML distance model 106c and the prediction model 106d are configured based on the output selected parameter set 310. The DL toxicity model of FIG. 3b includes the trained ML LD embedding model 106b as described with reference to FIG. 3b and the trained ML distance model 106c as described with reference to FIG. 3b. The toxicity prediction system 106 includes the ML phenotypic model 106a as described with reference to FIGS. 1a to 2 and the DL toxicity model of FIG. 3b for predicting the toxicity of one or more compounds applied to a plurality of test samples of a cell structure in an in vitro microscopy assay. The plurality of test samples in the sample test well group 306 of FIG. 3a can be captured by a microscopy imager during the in vitro microscopy assay. During the in vitro microscopy assay, a set of one or more compounds can be applied to the test samples, i.e., one compound per test sample. These can be processed and / or enhanced using the imaging system 104 of FIG. 1a or the quality control system and process of FIGS. 5a to 5f. The resulting images of the test samples can then be input into the toxicity prediction system 106 for predicting the toxicity of the compounds applied to the test samples. The toxicity prediction process 320 includes the following steps:
[0086] In step 321, an image set associated with one or more test samples to which a compound has been applied is received from the sample test well group 306 determined by in vitro microscopy. The compound may have known or unknown toxicity to the cell structure within the corresponding test sample. Each image of the test sample includes image data that sufficiently describes the cell structure of the relevant test sample to which the compound has been applied for automated processing and analysis. In step 322, each image in the image set is input into the trained ML phenotypic feature extraction model 106a or 200 of FIG. 1a or 2. For each input image of the test sample, the trained ML phenotypic feature extraction model 106a or 200 outputs a high-dimensional phenotypic representation of the cell structure of the test sample located within the input image of the test sample. Then, the output high-dimensional phenotypic representation of each test sample is applied to the trained ML LD embedding model 106b. In step 323, each high-dimensional phenotypic representation in the high-dimensional phenotypic representation of each test sample is input into the trained ML LD embedding model 106b for outputting a low-dimensional phenotypic embedding of each test sample. In step 324, the LD phenotypic embedding of each test sample is passed to the trained ML distance model 106c.
[0087] In step 325, the ML distance model 106c receives each LD phenotypic embedding 108 of each test sample and outputs a distance estimate 109 between the LD phenotypic embedding of each test sample and the set of NC LD phenotypic embeddings. In step 325a, the ML distance model 106c may be configured to output a distance or similarity estimate 109 between the LD phenotypic embedding of each test sample and the average NC LD embedding of the set of NC LD phenotypic embeddings. This may include, in step 325b, comparing the distance between the LD phenotypic embedding of the test sample and the average NC LD embedding of the set of NC LD phenotypic embeddings. From step 325, the ML distance / prediction model 106c / 106d may output an indication or probability associated with the distance of each LD phenotypic embedding in the LD phenotypic embedding to the NC LD embedding (i.e., determining how far the LD phenotypic embedding is from the non-toxic NC LD phenotypic embedding). In step 326, based on the comparison of the ML distance model 106c, an indication (e.g., probability) of the toxicity of each test sample and the compound applied to the test sample is output for each test sample.
[0088] FIG. 4a is a schematic diagram showing another example assay plate 400 having a group 402 of negative control samples and a group 404 of positive control samples for training a deep learning (DL) toxicity model of the toxicity prediction system 106a and the models described with reference to FIGS. 1a, 2, and 3a - 3b, and a group 406 of test samples for input into the trained DL toxicity model of the toxicity prediction system 106.
[0089] In this example, the HepaRG hepatocyte cell line is used to determine the toxicity of compounds in the liver. Each sample well in the sample plate 400 is filled with HepaRG cell constructs, wherein the in vitro microassay system 102 of FIG. 1a is configured to use a specific fluorescent substrate (e.g., carboxy-DCFDA (5-(and-6)-carboxy-2',7'-dichlorofluorescein diacetate) or CDFDA) to evaluate the cellular cholestatic effect of compounds. CDFDA is a reagent that diffuses passively into cells. It is a fluorescent substrate for the imager to capture the accumulation of CDF in the bile canaliculi in the image of each sample. This enables the evaluation of whether a compound causes cholestasis, which has been observed to occur when the bile canaliculi of the cell constructs disappear in the image. However, the toxicity prediction system 106 performs further unbiased processing to take into account other unobserved changes in the cell constructs when determining compound toxicity.
[0090] Samples with HepaRG cell constructs in the negative control sample group 402 have only buffer (e.g., DMSO) with no toxic effect applied to these samples. Samples with HepaRG cell constructs in the positive control sample group 404 have a reference compound with a known toxic effect (e.g., CPZ at 60 micromoles), which is known to be toxic to hepatocytes and trigger cholestasis. Samples with HepaRG cell constructs in the test sample group 406 have various different compounds at different concentrations and with repeated applications. In this example, each compound can be represented by 8 doses in the assay plate 400, wherein each dose is represented by 3 replicates in the assay plate 400. This is to ensure there are sufficient replicates such that after applying quality control to the images of the test samples, there will be at least one replicate for each compound and each dose with a viable image of the test sample for further downstream analysis.
[0091] Images of each sample well in the sample wells of the capture assay plate 400 are captured and the feasibility related to further downstream analysis is evaluated by the toxicity prediction system 106. This can be performed by the image system 104 of FIG. 1a and the image quality control system and / or process of FIGS. 5a-5e. For example, a quality control model can be trained to automatically evaluate the feasibility of each well sample by classifying the relevant images of the test samples in the wells. Each well sample is classified with a probability indicating whether the well sample is of good quality or poor quality. The higher the probability value, the better the quality, and the lower the probability value, the worse the quality. A threshold probability can be used to determine the viable higher quality sample wells. In this example, a probability value of approximately 0.1 is determined to yield viable sample images that can be passed to the toxicity prediction system 106 for toxicity training and / or toxicity prediction. The lightly shaded areas (e.g., images of the samples in wells 402a-402c, 404a, 406a-406c) represent samples in the NC, PC, and test sample wells indicating viable samples for further downstream analysis. The darkest shaded areas (e.g., images of the samples in wells 402d, 406b, and 404d) represent samples in the NC, PC, and test sample wells indicating non-viable samples with too many artifacts for analysis and that can be discarded. Thus, the set of viable sample images from each of the NC group 402, PC group 404, and test sample group 406 can be used for further downstream analysis. The set of viable images from the NC sample group 402 and the PC sample group 404 is used to train the DL toxicity model in an unsupervised manner, which includes training the ML LD embedding model 106b and the ML distance model 106c / 106d of FIG. 1a, as described with reference to FIGS. 1a-3b. The toxicity prediction system includes the ML phenotypic feature extraction model 106a and the DL toxicity model, which includes the ML LD embedding model 106b and the ML distance model 106c / 106d of FIG. 1a.
[0092] FIG. 4b shows an example distance matrix of negative and positive control samples of a trained DL toxicity model trained based on the unsupervised training process of FIGS. 1a and 3a using viable NC and PC samples from the NC sample group 402 and the PC sample group 406 on the assay plate 400 of FIG. 4a. Once the viable images of the NC and PC samples have passed through the ML phenotypic feature extraction model 106a or 200 of FIGS. 1a-2, the output set of the NC and PC high-dimensional phenotypic embeddings is used as the input to the UMAP algorithm to train and generate the ML LD embedding model 106b, as described with reference to FIG. 3b.
[0093] For example, a grid search is used to iteratively optimize a parameter set that defines hyperparameters of the UMAP algorithm, where different hyperparameter values from the parameter set are iteratively applied to determine the best combination of UMAP parameters that maximizes the difference between the NC low-dimensional phenotypic embedding set and the PC low-dimensional embedding set as much as possible. Assuming that the LD phenotypic embedding remains high-dimensional (e.g., 64 elements), Euclidean, Manhattan, and other standard distance metrics cannot be applied, and thus the Wasserstein distance is used instead.
[0094] As described with reference to FIGS. 1a and 3b, the output sets of the NC and PC LD phenotypic embeddings from the ML LD embedding model 106b based on the UMAP algorithm are applied to the Sinkhorn algorithm for training and generating the ML distance model 106c, as described with reference to FIG. 3b. The Wasserstein distance is optimized by iteratively performing a grid search within this parameter set to find the hyperparameters of the Sinkhorn algorithm that maximize the estimated Wasserstein distance between this NC LD phenotypic embedding set and this PC LD phenotypic embedding set. Then, the optimized hyperparameters output by the Sinkhorn algorithm can be used by the ML distance model 106c for other test sample LD embeddings to enable comparison of the distances with the negative control LD embeddings.
[0095] The matrix distance plot in FIG. 4b is a mapping of the distances when the toxicity prediction model 106 has been calibrated to use the viable input sample images from the assay plate 400 in FIG. 4a. Columns 0 to 9 and rows 0 to 9 of the matrix plot represent the distances between the NC LD embeddings associated with the viable image samples of the NC sample wells 402 in FIG. 4a. Columns 10-19 and rows 10-19 of the matrix plot represent the distances between the PC LD embeddings associated with the viable image samples of the PC sample wells 404 in FIG. 4a. Clearly, the DL toxicity model (i.e., the ML LD embedding model 106b and the ML distance model 106c / d) has been trained such that the distances between the NC LD embeddings have a dark gray shaded region 412 indicating the minimum distance therebetween. Similarly, for the PC LD embeddings, it also has a dark gray shaded region 414 indicating the minimum distance therebetween. Also, for this example, the light gray shaded region 416 indicates that the distance between this NC LD embedding set and this PC LD embedding set has been maximized (as much as possible). Assuming that the phenotypic distances in region 416 are maximized, it indicates that the corresponding PC samples have a toxic effect on the samples of the HepaRG cell line. This clearly indicates that the DL toxicity model of the toxicity prediction system 106 has been calibrated and can be used to test the toxicity of a series of compounds related to the HepaRG cell line.
[0096] Figure 4c shows the negative and positive control samples for predicting the toxicity of the compounds of the test samples using the trained DL toxicity models 106b - 106d of Figure 4b of the toxicity prediction system 106, as well as another example distance matrix 420 of the test samples. In addition to the viable NC and PC sample images output from the in vitro microassay represented by the assay plate 400 of Figure 4a, multiple viable test sample images are processed to predict the toxicity of the compounds applied to the test samples. In matrix 420, columns 0 to 95 and 288 to 311, and rows 0 to 95 and 288 - 311 represent the viable images of the NC samples used from plate 400, columns 312 to 335 and rows 312 - 335 represent the viable images of the PC samples used from plate 400 to train the DL toxicity models 106b - 106d of Figure 4b, and columns 96 to 287 represent the viable images of the test samples used from plate 400 for testing. As can be seen, the dark gray area 422 in columns 0 to 95 and rows 0 to 95 of the distance matrix represents the distances between the NC LD embeddings, which have the minimum distance between them. There are some false positives given a light gray shade, which do not affect the performance of the toxicity prediction. Similarly, the dark gray area 424 in columns 312 to 335 and rows 312 to 335 of the distance matrix 420 represents the distances between the PC LD embeddings, which also have the minimum distance between them. Again, the light gray shaded area 426 of the distance matrix indicates the distance between the set of NC LD embeddings and the set of test sample LD embeddings, which does not indicate the minimum distance from the NC LD embeddings, but rather a larger distance that has a higher toxic effect on the HepaRG cell line samples used in most, if not all, of the test samples.
[0097] Figure 4d shows an example of a conventional toxicity prediction method used for a set of 14 compounds (e.g., compounds A, B, C, D, E, F, G, H, I, J, K, L, M, and N) applied to the HepaRG cell structure. In this example, the assay plate has 14 compounds at 8 doses, with 3 replicates for each dose. 8 of these compounds are known to have toxic effects on the HepaRG cell structure and thus on the liver. These compounds are tested using commercial software accompanied by in vitro microscopy hardware and conventional toxicity analysis. This conventional toxicity analysis is based on using standard image analysis tools to analyze the data (e.g., stage 1), and a standard toxicity characterization based on vesicle counting (e.g., stage 2), as well as the researchers' analysis of the dose-response (EC50) graph (e.g., stage 3). In stage 4, it was found that only 8 compounds (e.g., compounds A, E, F, I, J, K, M, and N) had toxic effects, but the conventional toxicity analysis method determined that 6 toxic compounds (e.g., compounds B, C, D, G, H, and L) had no toxic effects. However, it is clear that the conventional toxicity analysis workflow did indeed miss compounds that had toxic effects due to the inability to detect 6 compounds. When these compounds were applied to the toxicity prediction system 106 and the trained DL toxicity models 106b / 106c, as described with reference to Figures 4a - 4c, it was found that all compounds had toxic effects. Figure 4e is a schematic diagram showing an example of the conventional toxicity prediction results of compounds A, B, C, D, E, F, G, and I obtained using the conventional method for predicting toxic compounds. The bile vesicle count and cell count of the positive control compound (i.e., chlorpromazine) are indicated within the dashed box, and the bile vesicle count and cell count of the negative control compound (i.e., DMSO) are indicated by the solid box. As can be seen, it is difficult to predict whether compounds B, C, D, and G have toxic effects when only considering the bile vesicle count and cell count. Figure 4f is a schematic diagram showing an example of the toxicity prediction results of the trained DL toxicity models of compounds A, B, C, D, E, F, G, and I obtained using the trained DL toxicity model 106b / 106c. The schematic diagram of Figure 4f is a graphical representation of a novel phenotypic distance metric (y-axis) of deep learning-based phenotypes extracted using a cell-based imaging classification backbone (e.g., Resnet50). The phenotypic distance metric of the positive control compound (i.e., chlorpromazine) (dashed box) and the negative control compound (i.e., DMSO (in the solid box)) is shown. The phenotypic distance metric compares each compound to the negative control compound (average phenotypic representation), so the distance (y-axis) between the tested compounds (e.g., compounds A, B, C, D, E, F, G, and I) and the average negative DMSO phenotypic profile is greater. As can be seen, for each tested compound, the distance between the compound phenotype and the reference negative profile (average DMSO) is above the 75th percentile of the DMSO distribution.In addition to the similarity with the positive control compound (chlorpromazine), this difference can also characterize the phenotypic effect as toxicity (the difference from DMSO alone cannot be characterized as toxicity alone). As can be seen, using the trained DL toxicity models 106b / 106c, compounds A, B, C, D, E, F, G, and I are more likely to be predicted as toxic, and compared with conventional toxicity prediction methods, the results can be used to more effectively predict the toxicity of compounds.
[0098] Figure 5a shows an example of a quality control process 500 in the image analysis and feature extraction system 104 of the toxicity pipeline 100 for Figure 1a, which is used to automatically identify viable and analyzable images of samples 103 captured from the HTS in vitro microscopy system 102 of Figure 1a. The QC process 500 is used to determine which samples or groups of samples from this set of wells on the assay plate are analyzable and non - analyzable. The QC process 500 uses ML techniques to help determine which well images captures can be retained or discarded for downstream analysis. This set of viable images is analyzable because the cell structure is well - defined enough for downstream image analysis processing, such as but not limited to, extracting phenotypic feature representations from the cell structure of the images for training and / or input into the DL / ML models 106b - 106c of the toxicity prediction system 106, as described with reference to Figures 1a to 4e. For example, this set of viable images 105 output from the QC process 500 improves the robustness and reliability of the trained DL models 106a - 106c, which further enhances the robustness and accuracy of the toxicity prediction output by the toxicity prediction system 106 relative to the toxicity prediction of the test samples to which the compounds are applied.
[0099] Although this set of viable images 105 output by the QC process 500 is described with reference to the toxicity prediction system 106, this is only an example, and those skilled in the art should understand that this set of viable images 105 output by the QC process 500 can be input into any other downstream process associated with analyzing, classifying, and / or predicting / estimating one or more aspects or characteristics of the images, depending on the type of HTS in vitro microscopy assay performed. For example, the HTS in vitro microscopy assay can be used in a drug / compound research program to determine the properties or effects of one or more compounds on cell structure samples of the in vitro microscopy assay, such as but not limited to, the toxicity of the compound applied to the cell structure sample, the efficacy effect of the compound on the cell structure sample, and / or any other type of property of the compound when applied to the cell structure sample in the in vitro microscopy assay, etc.
[0100] In this example, the QC process 500 identifies a first set of viable images of the sample from the images 103 captured by the HTS in vitro microscopy system 102 and preprocesses the first set of viable images into a final set of viable images 105 for input into (but not limited to) for example a toxicity prediction system 106 and / or any other compound analysis system or downstream workflow process / analysis system. Referring to the toxicity prediction system 106, the QC process 500 outputs the set of viable images of the sample for training the DL toxicity models 106a / 106b / 106c of the toxicity prediction system 106 of FIG. 1a and / or as an input to the trained DL models 106a / 106b / 106c of the toxicity prediction system 106, as described with reference to FIGS. 1a to 4e. The toxicity prediction system 106 receives the set of viable images associated with the plurality of samples, wherein at least one set of the plurality of samples has a compound applied thereto for testing and / or serves as a positive control group for training the DL models 106b-106c of the toxicity prediction system 106.
[0101] Upon receiving a set of images captured from a plurality of samples of an assay plate (e.g., assay plates 102c, 300, or 400 of FIGS. 1a, 3a, or 4a), the QC process 500 identifies viable samples of cell structures from the in vitro microscopy assay for further downstream analysis. The QC process 500 includes the following steps:
[0102] In step 502, a first set of sample images available for analysis is automatically identified from the received set of sample images captured from the plurality of samples of the assay plate. For example, the received set of sample images is automatically analyzed to determine whether features of the spheroids / cell structures of the sample are present and to determine whether these features highlight possible artifacts and / or other imaging defects outside the image or the focused region. Artifacts can bias predictions relative to downstream analysis using (a) ML model(s) etc. For example, the toxicity prediction system 106 uses the images of the received samples 105 to predict whether a compound is toxic or non-toxic, so any artifacts can be harmful to the training and / or prediction of toxicity as they may mask the effect of the compound applied to the cell structure. For example, these artifacts can mask the dose at which the compound is actually toxic (e.g., the dose with a 50% toxic effect).
[0103] Automated identification uses one or more ML models to estimate and predict whether the samples in each well are analyzable. For example, automatically identifying a first set of sample images can include inputting the received set of sample images from an assay into one or more ML models that are configured to identify one or more regions in each image of the set that may have cellular structures therein. The set of images with the identified regions can be input into one or more additional ML models to identify whether the cellular structures within the identified regions of each image are suitable for further downstream analysis. Those images from the received set of images that have regions containing cellular structures determined to be analyzable are selected to form the first set of sample images that are viable for further downstream analysis.
[0104] In step 504, for each image of the samples in the first set of sample images, a 2-dimensional (2D) set of images of the sample is generated, where the 2D set of images of the sample includes a plurality of 2D image slices of the sample captured in an assay plate, and the plurality of 2D image slices are taken along the z-axis of the sample. The 2D set of images of each sample in the first set of sample images is taken at different z-axis positions such that a 3D representation of each sample is formed. The plurality of 2D images of each sample form a 3D representation of each well of each sample, and the 3D representation can be further processed and represented as a fused / compressed 2D representation.
[0105] In step 506, a set of viable samples is identified from the set of 2D image slices. For example, for each sample in the first set of sample images, the set of 2D image slices can be fused, compressed, or combined to form a 2D representation of the 3D structure of the sample, where the 2D representation enhances the cellular structures within the sample. The cellular structures of each fused 2D representation can be further analyzed to identify whether they are sufficiently distinguishable or well-defined for downstream analysis (e.g., for input into the toxicity prediction system 106). For example, a fused 2D representation is sufficiently distinguishable or well-defined when it is determined that the quality of the image is suitable for extracting one or more metrics or phenotypic representations of the cellular structures from the image. Those samples from the first set of sample images with 2D image slice sets that are identified as sufficiently distinguishable or suitable for further analysis form the set of viable samples for analysis.
[0106] In step 508, data representing the set of viable samples for analysis is output as the image set 105. For example, each fused 2D representation image of the set of viable samples can form the image set 105 of FIG. 1a, and the image set can be input into a downstream process for analysis, such as the toxicity prediction system 106.
[0107] Figure 5b shows an example of a first quality control (QC) process 510 in step 502 of the QC process 500 for Figure 5a. The first QC process is for selecting a first viable image set of samples suitable for further downstream analysis. The first QC process 510 is configured to automatically identify a first set of sample images that are viable for analysis, e.g., for input into a toxicity prediction system 106 as described with reference to Figures 1a to 5a. The first QC process 510 may include the following steps: In step 512, each image of the samples of the received sample image set 103 from the HTS in vitro assay system 102 is preprocessed to form a preprocessed image set of the samples. The preprocessing may include performing image analysis and / or processing based on, but not limited to, for example, correcting illumination, lighting, illumination field correction, artifact reduction / interpolation, or any other image processing algorithm for processing the images to further enhance the cell structures that may be included in the images. In step 514, the preprocessed image set of the samples is input into a machine learning (ML) region of interest (ROI) model that is trained and configured to identify, for each input image of the sample, one or more regions of interest in the input image of the sample, where the regions of interest include the cell structures of the corresponding sample. The ML region of interest (ROI) model may be based on any type of neural network (NN) structure, such as, but not limited to, for example, a feedforward NN (FNN), a recurrent NN (RNN), an artificial NN (ANN), a convolutional NN (CNN), any other type of NN suitable for identifying regions of interest associated with cell structures in an image, a modification thereof, a combination thereof. For example, the ML ROI model may be based on a CNN that is trained to identify ROIs associated with cell structures in the input image and / or an output image that focuses on the identified ROIs associated with cell structures in the input image. The ML ROI model may be trained using a CNN structure and a labeled or annotated training image data set, where each image in the training image data set is labeled or annotated with information associated with whether there are cell structures and, if there are cell structures, the regions of interest within the image.
[0108] At this stage, those images in the input sample image set that the ML ROI model does not detect regions of interest in can be discarded from the sample image set, as these images are more likely to be non - analyzable. Alternatively, for each image in which the ML ROI model does not detect regions of interest, the region of interest may default to the entire input image for further analysis in step 516. That is, the region of interest can be identified as the entire input image, and the entire input image can be input into step 516.
[0109] In step 516, each image in the image of the identified region of interest having a cellular structure containing a sample is input into an ML image feasibility model that is trained and configured to classify whether the cellular structure in the input image is analyzable. Thus, each input image can be classified with a label or probability value representing the feasibility of the analyzable input image. If the label or probability value indicates that the input image is analyzable (e.g., the probability value can be greater than or equal to a predetermined feasibility probability threshold, or the label indicates that the image is feasible), the input image is placed into a first set of sample images that are considered to be feasible. The remaining input images having a label indicating infeasibility or a probability value less than the predetermined feasibility probability threshold are discarded.
[0110] The image feasibility model can be based on, but not limited to, for example, any ML algorithm or technique that can be used to classify whether the cellular structure within the ROI in the image is analyzable. This can be based on identifying whether any phenotypic features associated with the cellular structure (e.g., spheroids, cell nuclei, vesicles, etc.) within the input image exist in each input image, and evaluating whether there are sufficient phenotypic features present in the cellular structure of the sample captured in the image to enable further analysis of the cellular structure of the sample within the image. For example, the ML algorithm used to train the image feasibility model can be based on, but not limited to, for example, a support vector machine (SVM) or an NN classifier / structure, etc. As an example, the image feasibility model can be based on an SVM ensemble, where each SVM is associated with classifying a specific phenotypic feature of the cellular structure of the sample captured in the input image or classifying a cluster of phenotypic features, and where each SVM outputs a positive or negative classification indicating whether a specific phenotypic feature exists. If a sufficient number of SVM outputs indicate a positive classification of the corresponding specific phenotypic feature, the input image can be considered an image that is feasible for analysis.
[0111] In step 518, a first set of sample images is output that includes data representing those sample images that were classified as analyzable in step 516. The output first set of sample images can be further processed, for example but not limited to, in steps 504 - 508 of the QC process 500 to further enhance the cellular structure contained in each image and / or for further selection related to the feasibility of the first set of sample images. Alternatively or as an option, the first set of sample images that are considered to be feasible output from the first QC process 510 can be output from the image analysis system 104 as the image set 105 for input into the toxicity prediction system 106 or other downstream processes for analysis, etc.
[0112] FIG. 5c shows an example of an image analyzer 520 for implementing steps 512 and 514 of the first QC process 510 of FIG. 5b, where each image of the sample is preprocessed and regions of interest containing cell structures within each image of the sample are identified. In this example, each image in the set of sample images 103 output from the HTS in vitro microscopy assay system 102 is preprocessed by an image preprocessing unit 522. In this example, each input image 522a to the image preprocessing unit 522 has an illumination function 522b applied thereto, and the resulting illumination-corrected image 522c output from the first image preprocessing unit 522 is passed to an image ROI unit 524. The illumination function 522b may be configured to correct for non-uniform fluorescent illumination of the image capture device in the HTS in vitro microscopy system 102. Although the image preprocessing unit 522 performs illumination field correction on each image, this is merely an example and the present invention is not limited thereto. Those skilled in the art should understand that the image preprocessing unit 522 may perform one or more other image processing functions, such as image focusing, illumination field correction, sharpening, saturation, artifact reduction, correction and / or interpolation, any other image processing algorithms, combinations thereof, modifications thereto, etc., where the cell structures of the sample in each captured input image are further enhanced for use by the image ROI unit 524.
[0113] Each image in the pre - processed image set of the samples output from the image pre - processing unit 522 is input to the image ROI unit 524. The image ROI unit 524 is configured to include an ML region of interest (ROI) model based on a CNN architecture (e.g., segmentation based on CNN (CellPose(RTM))), which is trained and configured to identify each segment of the image and determine which segments contain cell structures, where the segments containing cell structures form the ROI of the image. The CNN - based segmentation (e.g., CellPose) architecture can be used on microscopy images to segment cell bodies, membranes, and nuclei. For example, CellPose is a deep - learning CNN architecture that has been trained on a highly variable cell image dataset containing over 70,000 segmented objects. These can be retrained and / or fine - tuned on an image training dataset focused on cell structures in microscopy assay images intended for downstream analysis. Although the CNN - based Cell Pose segmentation architecture is described, this is merely an example and the present invention is not limited thereto. Those skilled in the art should understand that any other suitable ML algorithm / model can be used, which performs the identification of the ROI containing cell structures related to the input image of the sample. The ML ROI model can be trained using a labeled training dataset including multiple images, each of which is annotated with a label or annotation that includes data indicating the presence or absence of a region of interest in the cell and / or the location of the region of interest within the image. For example, the training image dataset can include images labeled and annotated with information associated with the presence or absence of cell structures (e.g., phenotypic features of cell structures and / or macroscopic cell structures, cell spheroids, cell nuclei, cell mutations, perturbations, cell cholestasis, related biological features, etc.), and if there is a cell structure, a perturbed cell structure, or a remaining part of a cell structure, the region of interest contains the cell structure within the image. Thus, the ML ROI model is trained and configured to identify / predict the segments / ROI of the remaining / mutated / perturbed cell structures in the image of the sample in the input image and output this information for downstream processing.
[0114] For example, the input image of sample 524a may include cell structures, such as cell spheroid objects containing cell nuclei, where only those segments of the image containing cell nuclei are identified and form the region of interest (ROI) of the image. In another example, the input image of sample 524b may include cell spheroid objects representing cell structures undergoing cholestasis, i.e., the cell structures have been perturbed / mutated by a compound, and, but not limited to, for example, the cell nucleus may disappear, etc. In such a case, the ML ROI model has been trained to identify the segments / ROIs of the image containing the remaining / mutated / perturbed cell structures in the image of the sample. The image ROI unit 524 may be configured to output only those portions of image 524a that contain the ROI. Alternatively or additionally, the image ROI unit 524 may be configured to input image 524a, which is annotated with data representing the ROI, for use by other processing units / downstream processing when focusing on the cell structures contained in the ROI of the image. In any case, the image ROI unit 524 outputs an image data set that represents each pre-processed input image set in the pre-processed input image set of the sample, and for each image in the set, the ROI contains the cell structures of the sample within each input image. The output image and ROI data of the image analyzer 520 may be used in steps 516 and 518 of the first QC process 500, which may be implemented as an image feasibility system configured to extract features in the ROI of each image of the sample and determine whether each image is of good quality or poor quality, i.e., feasible or infeasible for further analysis, or whether each image is analyzable or non-analyzable. The output of the image feasibility system will include a set of images of analyzable / feasible samples and may form the image set 105, which may be input into further downstream processes, such as, but not limited to, for example, the toxicity prediction system 106.
[0115] FIG. 5d shows an example of an image feasibility system 530 as part of a QC pipeline that implements steps 516 and 518 of the first QC process 510 of FIG. 5b to classify whether the cellular structures in the ROI of each input image of a sample are analyzable (e.g., are feasible for further downstream analysis / processing). The image feasibility system 530 can include the image analyzer 520 of FIG. 5c that receives the set of microscopy images of the sample 103 output from the HTS microscopy assay system 102 of FIG. 1a and, for each image in the set of images 103, outputs data representing each image and an ROI containing the cellular structure features of each image. The set of images and the corresponding ROIs are input to a phenotype sampler 532 that, for each image in the set of images, identifies and / or classifies whether the ROI of each image has n phenotypic features (e.g., P1, P2, P3, …, Pn). Each of the n phenotypic features detected within the ROI of each image of the sample is clustered. Each of the n phenotypic features has a class of SVMs 536a - 536n that are trained to classify the ROI of an image with respect to one of the n phenotypic feature clusters. Each of the SVMs 536a - 536n outputs a classification of whether the ROI of the image has the corresponding phenotypic feature P1, P2, … Pn. Each of the n clusters of the n phenotypic features P1, P2, P3, …, Pn is input to the corresponding SVM 536a - 536n, and each SVM outputs the probability or likelihood that the corresponding phenotypic feature is present within the ROI of the image of the sample. The n output probabilities are combined to provide an output classification value 538 that indicates whether the image is analyzable (e.g., is feasible). Those images for which the classification value 538 indicates that the image is analyzable or feasible (e.g., the classification value 538 is greater than a feasibility threshold) are selected to form a set of feasible first images of the sample.
[0116] Figure 5e shows an example of a second QC process 540 in steps 504 - 508 of a quality control (QC) process 500 that enhances and selects a final set of viable images of samples for downstream analysis, such as but not limited to, for example, the input image set 105 of the toxicity prediction system 106 of FIG. 1a. Once a viable first set of images of the samples has been output from step 502 of the QC process 500 of FIG. 5a or from the image viability system 530 of FIG. 5c, step 504 of the QC process 500 can be performed, where, for each image of the samples from the first viable image set, a 3D representation of each well sample corresponding to each image of the samples is captured by a microscopy imaging device. The 3D representation of each image of the samples is generated by a microscopy imaging device (e.g., ImageXpress(RTM)) that captures a 2D image set of the samples, where the 2D image set of the samples includes a plurality of 2D image slices of the samples in the well samples of the assay plate. The plurality of 2D image slices are taken at different z - foci or different foci along the longitudinal z - axis of the sample well. The 2D image set of each image of the samples in the first viable image set is taken at different z - axis positions such that a 3D representation of each of the samples is formed. The 2D image set of each sample forms a 3D representation of the corresponding well for the image of the samples from the viable image set. Each 2D image set of each sample is enhanced by fusing or compressing / combining the 2D image set into a single 2D representation of the image of the sample to emphasize the cell structure in each 2D image. The second QC process 540 for enhancing and selecting a final set of viable images of samples for downstream analysis includes the following steps:
[0117] In step 542, for each 2D image set of the samples from the viable image set of the samples, the foreground, background, and a plurality of uncertain feature regions or cell structure classes in each 2D image slice from the 2D image set are identified. The plurality of uncertain feature regions or classes includes a plurality of uncertain foreground feature regions or classes and a plurality of uncertain background feature regions or classes. For example, for each image in the 2D image set, each x / y pixel value is divided into a plurality of classes, including but not limited to, for example, a foreground class, a background class, and a plurality of uncertain foreground classes and a plurality of uncertain background classes.
[0118] In step 544, for each 2D image set of the sample, the identified foreground, background, and multiple uncertain feature regions or classes of the 2D image slices in the 2D image set of the sample are iteratively combined to generate a single 2D image of the cellular structure of the sample. For example, the iterative process takes each foreground class and mixes it with the uncertain foreground class to attempt to optimize the projection between the two classes during the iterative process by smoothing the projection in a local region. The iterative optimization process can have two criteria. The final projection needs to be locally smooth in terms of the z-projection in the local region around the foreground class, and the local intensity in the local region is the maximum possible intensity. The iterative combination can be based on an ML smoothing model that iteratively smooths the foreground class and the uncertain foreground class.
[0119] In step 546, for each single 2D image of the sample, based on the quality of the single 2D image of the sample, the iteratively generated single 2D image of the sample is selected for the final set of viable image samples. For example, the quality of a single 2D image sample can be evaluated based on whether there are any sharp deviations between adjacent foreground feature regions and adjacent uncertain foreground regions, etc., or whether there are any sharp deviations between adjacent background feature regions and adjacent uncertain background regions. For example, if there are too many "jumps" between adjacent foreground regions / classes and adjacent uncertain foreground regions / classes at the end of the iterative process, or if the resulting single 2D image still does not meet the criteria related to the image intensity in the local region of the z-projection, or if the iterative process does not converge within an appropriate error threshold, then the single 2D image is not viable and can be selected for discard. The remaining single 2D images for each sample can be used to form the final set of viable images.
[0120] In step 548, a final set of viable images is output that includes image data representing the single 2D image of each sample selected in step 546. The final set of viable images can form this image set 105 input to the downstream process, such as but not limited to, for example.
[0121] Figure 5f shows another example of a Smooth Manifold Extraction (SME) system 550 for implementing steps 504 - 508 of the QC process 500 of FIG. 5a and / or steps 542 - 548 of the second QC process 540 of FIG. 5e, the second QC process being for further enhancing and selecting a final set of viable images 567a - 567p for output from the image analyzer system 104. The SME system 550 includes a chain of processing units including a 3D image representation unit 552, a profiling unit 554, an FFT unit 556, pixel classification units 558 - 560 for foreground / background classification / clustering, an iterative SME unit 562, and a final image selection unit 564 and an output unit 566. The output unit 566 outputs the final set of viable images 567a - 567p, which can be used as an input to the image set 105 for further downstream analysis processes such as, but not limited to, for example, input to a toxicity prediction system 106.
[0122] The 3D image representation unit 552 is configured to perform step 504 of FIG. 5a, which receives a first set of viable images of samples that have been output from step 502 of the QC process 500 of FIG. 5a or from the image viability system 530 of FIG. 5c. The 3D image representation unit 552 captures a 3D representation of the samples corresponding to each viable image of the samples in each sample well 551. For example, a guided microscopy imaging device 553 (e.g., ImageXpress(RTM)) is directed to capture a 2 - dimensional (2D) image set 555 of the sample wells 551 corresponding to the viable images of the samples. The 2D image set 555 of the samples includes a plurality of 2D image slices 555a - 555k of the samples from the corresponding sample wells 551 of the assay plate. The sample well 551 has a longitudinal z - axis 551a, an x - axis 551b, and a y - axis 551c, where the x - axis 551b and the y - axis 551c are orthogonal to each other and also orthogonal to the longitudinal z - axis. The 2D image slices 555a - 555k are taken at different z - foci or different foci along the longitudinal z - axis 551a of the sample well 551. The 2D image set 555 of each viable image of the samples is taken at different z - axis positions such that a 3D representation of each sample is formed. The 2D image set 555 of each sample forms a 3D representation of each well 551 corresponding to the viable images of each sample from the received set of viable images.
[0123] The remaining processing units 554-566 of the SME system 550 are configured to process each 2D image set 555 of each sample to enhance / emphasize the cellular structure of the sample represented in each 2D image. This is performed by fusing or compressing / combining each 2D image set 555 of each sample into a single 2D representation 565 of the sample. Substantially, these remaining processing units 554-566 are configured to find, for each sample, the z focus of each pixel of each 2D image set 555 corresponding to the relevant structure of the cell / cellular structure of each pixel.
[0124] In the dissection unit 554, a profile of each 2D image slice 555a of the 2D image set 555 is extracted, where any (x,y) location corresponds to a profile of the focus value in the z direction 551 of the sample well 551, and the profile is composed of the direct intensity value in the case of a confocal image or the SML value of a wide-field epi-fluorescence image. Each 2D image slice 555a-555k (e.g., the profile) passing through a certain foreground signal contains a lower frequency component. The profile of each 2D image slice in the 2D image slices 55a-55k is passed to the FFT unit 556. For each 2D image set of the sample, the FFT unit 556 performs a fast Fourier transform (FFT) on each profile in the profiles of the 2D image slices within the 2D image set 555 of the sample to determine the power spectrum of each profile of the 2D image slices 555a-555k. The FFT profiles of the 2D image slices 555a-555k are passed to the pixel classification units 558-560 for foreground / background classification and clustering.
[0125] The pixel classification units 558 - 560 perform multi-class classification on the FFT profiles of the 2D image slices 555a - 555k using an ML classification algorithm / model, thereby classifying each (x,y) location or pixel of each profile of the 2D image slices into one of multiple labels associated with a foreground class 559a, a background class 559b, and / or multiple uncertain foreground / background classes 559c - 559l. The ML classification algorithm can be based on, but not limited to, for example, a multi-class k-means algorithm, a multi-class UMAP algorithm, and / or any other suitable multi-class classification algorithm or model capable of classifying the (x,y) location of each pixel of each FFT profile in the FFT profiles of the 2D image slices 555a - 555k of the 2D image set 555 into a label associated with the foreground class 559a, the background class 559b, or the multiple uncertain foreground / background classes 559c - 559l. The pixel classification unit 556 outputs to the SME unit 562 the labeled / annotated FFT profiles 561 of the 2D image slices 555a - 555k, where each (x,y) pixel of each 2D image slice has been marked with a classification associated with one of the foreground class 559a, the background class 559b, or the multiple uncertain foreground / background classes 559c - 559l.
[0126] The SME unit 562 is based on an ML SME model / algorithm 563 (or an iterative SME optimization process) that iteratively processes data representing the 2D image set of each sample based on the determined foreground, background, and uncertain foreground and background classes of the pixels of the 2D image slices of the 2D image set assigned to each sample. The ML SME model 563 uses an iterative SME algorithm. The ML SME model 563 is configured to perform an iterative SME optimization process that combines data representing the 2D image slices of the sample using a cost function in which the balance of local smoothness and proximity to the maximum focus value is minimized, thereby obtaining the final smoothed index map 563m of the sample. It should be noted that the first index map 563a is highly discontinuous at the start of the iterative SME optimization process, but should be smoothed when the iterations of the SME optimization process converge to the final smoothed index map 563m while retaining fine details on the foreground. The final smoothed index map 565m is sent to the extraction and selection unit 564.
[0127] For example, the SME algorithm 563 of the SME unit 562 examines foreground pixels mixed with an uncertain foreground class and attempts to optimize the projection of these types of foreground classes of the 2D image slices 555a - 555k using an iterative SME optimization process and fuse / compress them into a final single 2D image 565a (or a final index map 563m). The iterative SME optimization process smooths the projection in each iteration from the first iteration (e.g., iteration 001) to the final iteration (e.g., iteration 571) when the final error threshold has been reached / satisfied and / or the maximum number of iterations has been satisfied / reached. The iterative SME optimization process needs to meet two criteria. The first criterion is that the final projection (e.g., the final index map 563m or the final single 2D image) needs to be locally smooth in terms of the z - projection. The second criterion is that the local intensity is the maximum possible intensity (e.g., using a Gaussian energy function). These criteria are used in each iteration from the first iteration to the final iteration when fusing / compressing the 2D image slices together to form the final single 2D image. Once the iterative SME optimization process of the SME algorithm 563 has converged to the final index map 563m and / or the maximum number of iterations has been satisfied, the resulting final index map 563m is used to form the final single 2D image 565a. The resulting final index map 563m or the final single 2D image 565a is passed to the extraction and selection unit 564 to determine whether the final single 2D image 565a should be included in the final set of viable images.
[0128] In the extraction and selection unit 564, the final smoothed index map 563m output from the iterative SME algorithm 563 is analyzed, and a plurality of voxels corresponding to the index map 563m are extracted from the original stack to generate the final single 2D image 565a. The extraction and selection unit 564 using the selection unit 565b selects whether to include the final single 2D image 565a in the final set of viable images of the sample for sending to the output unit 566. The selection unit 565b determines whether the final single 2D image 565a or the index map 563m corresponding to the final single 2D image 565a has sufficient quality or sufficient smoothness for output from the extraction and selection unit 564. For example, the selection unit 565b determines whether the iterative SME algorithm 563 or the final single 2D image 565a meets two criteria. The first criterion is based on determining whether the projection of the final single 2D image 565a or the final index map 563m of the iterative SME algorithm 563 is locally smooth in terms of the z - projection. The second criterion is based on determining whether the local intensity is at the maximum possible intensity.
[0129] For example, the first criterion is based on checking the smoothness of each local pixel region of the final single 2D image 565a (or the final index map 563m). In this case, for each pixel local region of the final single 2D image 565a (or the final index map 563m), such as but not limited to, for example, each 5×5 pixel neighborhood in the final 2D image 565a (or the final index map 563m), the z-level in each local region (e.g., 5×5 pixel neighborhood) should be smooth in terms of the z-level, where when moving from one pixel to another pixel in the local region, there should be no deviation or "jump" in intensity (e.g., not directly jumping from intensity level 1 to 17). The second criterion is based on the intensity of the z-projection of each local pixel region of the final single 2D image 565a (or the final index map 563m). This second criterion is satisfied when, within each local region (e.g., 5×5 pixel neighborhood), the z-projection obtained over the local region (e.g., 5×5 pixel neighborhood) has an intensity value based on the local region (e.g., 5×5 pixel neighborhood) with the maximum z-intensity. In the selection unit 565b, if these two main criteria are not met, the selection unit 565b discards the final single 2D image 565a (or the final index map 563m) that does not form part of the final viable image set 567. Otherwise, the final single 2D image 565a (or the final index map 563m) of the sample is output by the extraction and selection unit 564 for sending to the output unit 566, and the output final single 2D image 565a is added as the final viable image 567a of the final viable image set 567.
[0130] Although the above first and second criteria are described herein, this is merely by way of example, and those skilled in the art will understand that modifications to these two criteria and / or other criteria can be applied to select / discard the final single 2D image 565a when sending it to the output unit 566. For example, these modifications and / or other criteria can be based on but not limited to, for example, the first criterion can be further modified to count the number of "jumps" in intensity deviation for determining whether to discard / select the final single 2D image 565a (e.g., if the number of intensity deviation "jumps" is below a minimum "jump" threshold / count, the final single 2D image 565a can be selected for output, otherwise it is discarded), and / or the selection unit 565b can further consider whether the iterative SME algorithm converges within a predetermined error threshold (e.g., if the convergence error of the iterative SME algorithm is less than or equal to the predetermined error threshold, the final single 2D image 565a can be selected for output, otherwise it is discarded), and / or any other criteria for evaluating whether the final single 2D image 565a of each sample has sufficient quality and / or can be analyzed by further downstream processes, etc.
[0131] The output unit 566 stores each of the output final single 2D images 565a selected for output by the extraction / selection unit 564 as a plurality of images that form a final viable image set of the samples. The output unit 566 can then send the final viable image set of the samples to further downstream processes, such as but not limited to, for example, as an input image set 105 to the toxicity prediction system 106.
[0132] Figure 6a Is a schematic diagram of a system / apparatus for performing the methods described herein. The illustrated system / apparatus is an example of a computing device. Those skilled in the art will understand that other types of computing devices / systems can alternatively be used to implement the methods described herein, such as a distributed computing system.
[0133] The apparatus (or system) 600 includes one or more processors 602. The one or more processors control the operation of the other components of the system / apparatus 600. For example, the one or more processors 602 can include a general-purpose processor. The one or more processors 602 can be a single-core device or a multi-core device. The one or more processors 602 can include a central processing unit (CPU) or a graphics processing unit (GPU). Alternatively, the one or more processors 602 can include dedicated processing hardware, such as a RISC processor or programmable hardware with embedded firmware. Multiple processors can be included.
[0134] The system / apparatus includes a working or volatile memory 604. The one or more processors can access the volatile memory 604 to process data and can control the storage of data in the memory. The volatile memory 604 can include any type of RAM, such as static RAM (SRAM), dynamic RAM (DRAM), or it can include flash memory, such as an SD card.
[0135] The system / apparatus includes a non-volatile memory 606. The non-volatile memory 606 stores an operation instruction set 608 for controlling the operation of the processor 602 in the form of computer-readable instructions. The non-volatile memory 606 can be any type of memory, such as read-only memory (ROM), flash memory, or magnetic drive memory.
[0136] One or more processors 602 are configured to execute operational instructions 608 to cause the system / apparatus to perform any of the methods described herein. The operational instructions 608 may include code related to the hardware components of the system / apparatus 600 (i.e., drivers), as well as code related to the basic operations of the system / apparatus 600. Generally, one or more processors 602 execute one or more of the operational instructions 608 permanently or semi-permanently stored in the non-volatile memory 606, using the volatile memory 604 to temporarily store data generated during the execution of the operational instructions 608.
[0137] Figure 6b FIG. is a schematic diagram of a system 610 for performing the methods described herein. The illustrated system 610 is an example of a computing device, system, and / or cloud computing system, etc. Those skilled in the art will understand that other types of computing devices / systems may alternatively be used to implement the methods described herein, such as distributed computing systems. The system 610 includes a receiver module / unit 612, a first ML model module / unit 614, a second ML model module / unit 616, a distance comparison module / unit 618, and an output module / unit 620, which may be connected together or communicate with each other as needed to implement the methods and / or apparatus / systems described herein.
[0138] For example, the receiver module 612 may be configured to receive a set of images associated with the plurality of samples. The first ML model module 614 may be configured to input each image in the set of images into a first ML model 106a, which is configured to predict phenotypic features 107 of cell structures within the sample associated with each such image. The second ML model module 616 may be configured to input each of the predicted phenotypic features 107 associated with each sample into a second ML model 106b, which is configured to predict a lower-dimensional phenotypic feature embedding 108 of each such sample. The distance comparison module 618 may be configured to compare the distance between the lower-dimensional phenotypic feature embedding 108 of each sample and the lower-dimensional phenotypic feature embedding of a sample of a compound with known toxicity, which may output data representing the comparison 109. The output module 620 may be configured to output, for each sample, an indication of the toxicity of each sample and the compound applied to the sample based on the comparison 109.
[0139] Embodiments of the methods described herein can be implemented in digital electronic circuitry, integrated circuit systems, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These can include computer program products (such as software stored on, for example, a disk, an optical disc, a memory, a programmable logic device) that include computer-readable instructions that, when executed by a computer, such as described with respect to Figure 6a or 6b, cause the computer to perform one or more of the methods described herein.
[0140] Any system features described herein can also be provided as method features, and vice versa. As used herein, means-plus-function features can alternatively be expressed in terms of their corresponding structures. Specifically, method aspects can be applied to system aspects, and vice versa.
[0141] Furthermore, any, some, and / or all features in one aspect can be applied in any suitable combination to any, some, and / or all features in any other aspect. It should also be understood that the specific combinations of the various features described and defined in any aspect of the present invention can be implemented and / or provided and / or used independently.
[0142] Although several embodiments have been shown and described, those skilled in the art should understand that changes can be made in these embodiments without departing from the principles of the disclosure, the scope of which is defined in the claims.
Claims
1. A computer-implemented method for predicting the toxicity of one or more compounds applied to a plurality of samples of a cell structure (102a) in an in vitro microscopy assay (102), the method comprising: include: receiving (112) a set of images (105) associated with the plurality of samples; inputting (114) each image in the image set (105) into a first ML model (106a), the first ML model being configured to predict a phenotypic characteristic (107) of a cellular structure (102a) within a sample associated with said each image; inputting (116) each of the predicted phenotypic features (107) associated with each sample into a second ML model (106b) configured to predict a lower dimensional phenotypic feature embedding (108) for each sample; comparing (118) the distance between the lower dimensional phenotypic feature embedding of each sample and the lower dimensional phenotypic feature embedding of the sample to which a compound with known toxicity is applied; as well as Based on the comparison, an indication of the toxicity of each sample and the compound applied to the sample is output (120) for each sample.
2. The computer-implemented method of claim 1, in, The first ML model (106a) is a neural network or convolutional neural network (CNN) model (200), which is trained using cell image training data for classification, and the predicted phenotypic features are embedded within a complete layer of the trained neural network or CNN model (200), and the method includes outputting the phenotypic feature embeddings (107) from the complete layer.
3. A computer-implemented method as claimed in claim 1 or 2, in, The second ML model (106b) is based on a uniform manifold approximation and projection UMAP algorithm or a t-SNE algorithm for dimensionality reduction of the phenotypic feature embedding (107) for the sample, wherein the phenotypic feature embedding (107) is mapped to a lower dimensional vector space for comparing the distance between the phenotypic feature embedding and the phenotypic feature embedding of samples of compounds with known toxicity.
4. The computer-implemented method of claim 3, further comprising: include: training (313) the second ML model (108b) using unsupervised training based on negative control samples and positive control samples in the plurality of samples on a UMAP technique to predict a toxicity distance metric associated with the samples to which a compound with known toxicity is applied; or Training the second ML model (106b) includes iteratively performing a grid search over the set of hyperparameters of the UMAP technique to select those hyperparameters that maximize the difference between the negative control samples and the positive control samples.
5. A computer-implemented method as claimed in claim 3 or 4, in, Indicative of toxicity embedded in this phenotypic signature of samples to which the compound was applied include: applying (323) the phenotypic feature embedding (107) of the sample to which the compound was applied to the second ML model (106b) to output a lower dimensional embedding (108) of the sample to which the compound was applied; and Based on comparing the distance between the lower dimensional embedding (108) and embeddings of one or more samples to which a compound with known toxicity was applied, an indication of toxicity of the samples to which the compound was applied is determined (325, 325a, 325b, 326).
6. A computer-implemented method as claimed in claim 3 or 4, in, The phenotypic characteristics indicative of embedded toxicity in a sample to which the compound was applied further include: applying (323) the phenotypic feature embedding (107) of the sample to which the compound was applied to the second ML model (106b) to output a lower dimensional embedding (108) of the sample to which the compound was applied; and The lower dimensional embedding (108) of the sample to which the compound is applied is applied (324, 325) to a third ML model (106c) trained to output an indication of a distance (109) between the lower dimensional embedding (108) and a set of lower dimensional embeddings associated with the negative control samples.
7. The computer-implemented method of claim 6, further comprising training the third ML model (106c) based on performing a grid search on a set of hyperparameters of a high-dimensional distance metric algorithm that maximizes the distance between the lower dimensional embeddings of the negative control samples and the positive control samples while minimizing the distance between the lower dimensional embeddings of the negative control samples or minimizing the distance between the lower dimensional embeddings of the positive control samples.
8. A computer-implemented method as claimed in any preceding claim, in, Comparison of distances used to indicate the toxicity of the phenotypic feature embedding of the sample was based on the Wasserstein distance metric.
9. A computer-implemented method as claimed in any preceding claim, in, Receiving the set of images associated with the plurality of samples (105) further comprises: Viable samples of cellular structures are identified (500) for analysis in an in vitro microscopy assay based on the following steps: automatically identifying (502) a first set of samples available for analysis from a plurality of samples on an assay plate; generating (504) a 2-dimensional 2D image set for each sample in the first sample set, the 2D image set of each sample comprising a plurality of 2D image slices taken along a z-axis of each sample; and identifying (506) a feasible sample set from the set of 2D image slices; and The output (508) represents data of the feasible sample set for analysis as the image set (105).
10. The computer-implemented method of claim 9, in, Automatically identifying (502) the first set of samples further comprises, for each sample in the plurality of samples: preprocessing (512) the image of each sample; Inputting (514) the pre-processed sample image to a first machine learning (ML) model, the first ML model being configured to identify a region of interest including a cell structure of the input sample image; inputting (516) the identified region of interest of the sample image to a second ML model configured to classify whether the sample is analyzable; and The output (518) includes the first sample set of data representing those samples classified as analyzable.
11. The computer-implemented method of claim 10, in, The first ML model (520) for identifying the region of interest is a convolutional neural network (CNN) or other neural network trained to identify regions of interest including cellular structures, and The second ML model (530) for classifying whether the sample is analyzable is a first-class SVM configured to classify whether the region of interest is analyzable. in, The CNN is trained / configured based on a labeled training dataset, wherein the labeled training dataset includes a plurality of images, each of the images being annotated with a label, the label including data indicating whether a cell region of interest exists and / or the location of the cell region of interest within the image; and The one-class SVM is trained / configured for classifying whether the region of interest is analyzable.
12. A computer-implemented method as claimed in any one of claims 9 to 11, in, Identifying the feasible sample set from the set of 2D image slices further comprises, for each sample: identifying (542) a foreground, a background, and a plurality of uncertain feature regions of the cell structure in each of the 2D image slices; iteratively combining (544) the foreground, background and the uncertain feature regions of the 2D image slices to generate a single 2D image of the cell structure; as well as selecting (546) the sample for the feasible sample set based on the quality of the single 2D image; and outputting, for each selected sample, a feasible image set, the feasible image set comprising a feasible image set representing image data of the single 2D images associated with each selected sample; The multiple uncertain feature regions include multiple uncertain foreground features and multiple uncertain background features.
13. An apparatus (600) comprising a processor (602), a memory unit (606) and a communication interface (604), in, The processor (602) is connected to the memory unit (606) and the communication interface (604), wherein the processor (602) and the memory (606) are configured to implement the computer-implemented method as claimed in any one of the preceding claims.
14. A computer readable medium comprising data or instruction codes which, when executed on a processor (602), cause the processor (602) to implement the computer-implemented method of any one of claims 1 to 12.
15. A system (610), include: a receiver module (612) configured to receive a set of images associated with the plurality of samples; a first ML model module (614) configured to input each image in the image set into a first ML model (106a), the first ML model configured to predict a phenotypic feature (107) of a cellular structure within a sample associated with each image; a second ML model module (616) configured to input each of the predicted phenotypic features (107) associated with each sample into a second ML model (106b) configured to predict a lower dimensional phenotypic feature embedding (108) for each sample; a distance comparison module (618) configured to compare the distance between the lower dimensional phenotypic feature embedding (108) of each sample and the lower dimensional phenotypic feature embedding of samples of compounds with known toxicity; as well as An output module (620) is configured to output, for each sample, an indication of the toxicity of the compound applied to the sample based on the comparison (109).