Quality control of in vitro analysis sample output

By automatically identifying and analyzing cellular structure samples in high-throughput screening, using machine learning models to identify regions of interest and generating high-quality 2D images, the problem of inaccurate identification of compounds in the prior art is solved, and reliable sample screening and drug discovery is achieved.

CN120051813APending Publication Date: 2025-05-27SANOFI SA(FR)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380073062.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-17
Filing Date
2023-10-13
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to reliably identify the non-toxicity or toxicity of compounds in high-throughput screening, resulting in an increased risk of drug-induced liver damage, etc.

Method used

Through a computer-implemented method, the sample sets that can be used for analysis on the assay board are automatically identified, the 2D image sets are generated, the feasible sample sets are identified, and the data representing the feasible sample sets are outputted for analysis. The method includes identifying regions of interest in cellular structures using machine learning models and generating a single 2D image by iteratively combining 2D image slices to select high-quality viable samples.

Benefits of technology

Efficient quality control of artifacts and error data points in HTS data is achieved, and a reliable and feasible sample set is generated, reducing the risk of compounds being tested in vivo and improving the reliability of drug discovery programs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120051813A_ABST
    Figure CN120051813A_ABST
Patent Text Reader

Abstract

Methods, devices, systems, and computer-implemented methods configured for identifying viable samples of cellular structures for analysis in in vitro microscopy assays are disclosed. A first set of samples available for analysis is automatically identified from a plurality of samples of an assay plate. A two-dimensional (2D) image set is generated for each sample in the first set of samples. The 2D image set for each sample includes a plurality of 2D image slices taken along the z-axis of each sample. A set of feasible samples is identified from the set of 2D image slices. Data representing the set of feasible samples for analysis is output as the set of images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present description relates to devices, systems, and method(s) for identifying viable samples of cellular structures for downstream analysis of in vitro microscopy assays. Background Art

[0002] Cell structures that can imitate and / or simulate the process and / or function of an organ of a subject or patient have been developed. These cell structures can be used for the in vitro test of the efficacy, non-toxicity and / or toxicity of various compounds related to the organ, rather than in vivo testing. This cell structure can include an immortalized cell line that has been developed to imitate or simulate a specific organ of a subject. This has realized a semi-automatic toxicity prediction test system and method, which can be used to use image microscopy and observe the changes caused by non-toxicity / toxicity through dose-response (DR) charts, etc. to identify compounds and / or the efficacy of non-toxic / toxic compounds to the organs of a subject.

[0003] Conventional semi-automatic test systems and methods can be used to identify compounds that produce detectable signals to assess the impact of compounds on the cell structures associated with organs. These test systems are referred to as assays. Once an assay for non-toxic / toxicity prediction is developed, researchers can use the assay to identify compounds with the desired activity associated with non-toxic / toxicity. Typically, the compound will be tested at multiple concentrations, image microscopy will be used for imaging, and DR graphs or other metrics useful for researchers to determine their non-toxic / toxicity can be generated. For example, the analysis of the DR chart can allow researchers to determine whether the compound is active, non-toxic and / or toxic, and at what concentration.

[0004] It is desirable to test large numbers of potential compounds, which is typically done using high throughput screening (HTS). This uses robotics, data processing / control and imaging software, liquid handling equipment, and sensitive detectors, and allows researchers to rapidly perform thousands or even millions of screening tests. However, the large amounts of data generated at the imaging and DR steps of HTS activities require careful analysis by researchers to detect artifacts and correct erroneous data points before validating the experiment.

[0005] Considering that a large amount of data is generated from HTS, and even researchers detect artifacts and correct erroneous data points through post-screening analysis, it has been found that even semi-automatic non-toxicity / toxicity prediction assays using HTS output data cannot reliably identify the non-toxicity or toxicity of each compound. This is especially true when the known compound is non-toxic or toxic when analyzed on the cellular structure of the simulated / mimicked organ. This increases the risk of performing in vivo tests on compounds that have passed such semi-automatic non-toxicity / toxicity prediction assays. For example, 20%-40% of patients with drug-induced liver injury (DILI) present with cholestasis and / or mixed hepatocellular / cholestatic injury patterns. Drug-induced hepatotoxicity or DILI is an acute or chronic reaction to natural or manufactured compounds. Conventionally, DILI can be classified based on clinical manifestations (hepatocellular, cholestatic or mixed), hepatotoxic mechanisms, or histological appearance from liver biopsy. Therefore, reliable compound in vitro non-toxicity / toxicity prediction is an important part of drug / compound discovery or research plans.

[0006] Improved methods, devices, systems and / or architectures that can perform quality control to efficiently and reliably detect artifacts / erroneous data points in HTS data and generate viable sample sets for downstream analysis (such as, but not limited to, for example, predicting the efficacy, non-toxicity or toxicity of compounds on cellular structures from samples output by in vitro HTS assays, etc.) are desired. Summary of the invention

[0007] According to a first aspect, a computer-implemented method for identifying viable samples of cellular structures for analysis in an in vitro microscopy assay is provided, the method comprising: automatically identifying a first set of samples available for analysis from a plurality of samples in an assay plate; generating a 2-dimensional (2D) image set for each sample in the first set of samples, the 2D image set of each sample comprising a plurality of 2D image slices captured along a z-axis of the each sample; identifying a viable sample set from the set of 2D image slices; and outputting data representing the viable sample set for analysis as the image set.

[0008] A computer-implemented method as described in the first aspect, wherein automatically identifying the first sample set further includes, for each sample of the multiple samples: preprocessing an image of each sample; inputting the preprocessed sample image into a first machine learning (ML) model, the first ML model being configured to identify a region of interest including a cell structure of the input sample image; inputting the identified region of interest of the sample image into a second ML model, the second ML model being configured to classify whether the sample is analyzable; and outputting the first sample set including data representing those samples classified as analyzable.

[0009] A computer-implemented method as described in the first aspect, wherein the first ML model is a convolutional neural network (CNN) or other neural network trained to identify a region of interest including a cell structure, and the second ML model is a class SVM configured to classify whether the region of interest is analyzable. As an option, the CNN is trained and configured based on a labeled training data set, wherein the labeled training data set includes a plurality of images, each of the images being annotated with a label, the label including data indicating whether a cell region of interest exists and / or the location of the region of interest within the image. As an option, the class SVM configured to classify whether the region of interest is analyzable is trained and configured.

[0010] A computer-implemented method as described in the first aspect, wherein identifying the feasible sample set from the 2D image slice set further includes, for each sample: identifying the foreground, background, and uncertain feature region of the cell structure in each of the 2D image slices; iteratively combining the foreground, background, and uncertain feature region of the 2D image slices to generate a single 2D image of the cell structure; and selecting the sample for the feasible sample set based on the quality of the single 2D image.

[0011] As an option, the plurality of uncertain feature regions include a plurality of uncertain foreground features and a plurality of uncertain background features.

[0012] A computer-implemented method as described in the first aspect, wherein identifying the feasible sample set from the set of 2D image slices further includes, for each sample: identifying the foreground, background, and multiple uncertain feature regions of the cell structure in each of the 2D image slices, wherein the multiple uncertain feature regions include multiple uncertain foreground features and multiple uncertain background features; iteratively combining the foreground, background, and multiple uncertain feature regions of the 2D image slices to generate a single 2D image of the cell structure; and selecting the sample for the feasible sample set based on the quality of the single 2D image; and outputting data representing an image of the feasible sample set associated with the feasible sample set.

[0013] In the computer-implemented method of the first aspect, outputting data representing the feasible sample set further includes outputting data representing an image of the feasible sample set. As an option, outputting data representing an image of the feasible sample set further includes outputting data representing one or more of the following: a set of 2D images generated for each feasible sample in the feasible sample set; a pre-processed image of each feasible sample in the feasible sample set; a single 2D image generated for each feasible sample, each single 2D image being generated by iteratively combining the 2D image slices of the feasible sample based on the identified foreground, background, and uncertainty regions of the 2D image slices of the feasible sample; and any other images captured or processed in relation to the feasible sample.

[0014] The computer-implemented method of the first aspect, wherein the cell structure comprises one or more from the group consisting of: a cell spheroid structure; a vesicle; an organoid; and any other suitable cell structure.

[0015] The computer-implemented method of the first aspect, wherein the in vitro microscopy assay is a high throughput screening in vitro microscopy assay. As an option, the plate comprises a plurality of wells, wherein each well has a sample of the cell structure therein.

[0016] According to the computer-implemented method of the first aspect, data representing each of these feasible samples is input into a third ML model, which is configured to perform downstream assay analysis on the feasible samples to predict the assay analysis results of each of these feasible samples.

[0017] The computer-implemented method of the first aspect, wherein the first subset of the samples comprises a negative control, the second subset of the samples comprises a positive control, and the third subset of the samples comprises samples to be analyzed, wherein the third ML model is trained based on the negative control / the positive control. As an option, the assay analysis comprises at least one item from the group consisting of: toxicity analysis; non-toxicity analysis; efficacy analysis; and any other analysis.

[0018] A computer-implemented method as described in the first aspect, wherein the assay analysis includes a toxicity analysis, which is configured to predict the toxicity of one or more compounds applied to multiple feasible samples of cell structures in the in vitro microscopy assay, the method comprising: receiving a set of images associated with the multiple samples; inputting each image in the image set into a first ML model, the first ML model being configured to predict a phenotypic feature of the cell structure within the sample associated with each image; inputting each of the predicted phenotypic features associated with each sample into a second ML model, the second ML model being configured to predict a lower dimensional phenotypic feature embedding for each sample; comparing the distance between the lower dimensional phenotypic feature embedding of each sample with the lower dimensional phenotypic feature embedding of a sample to which a compound with known toxicity is applied; and based on the comparison, outputting, for each sample, an indication of the toxicity of the compound applied to the sample.

[0019] A computer-implemented method as described in the first aspect, wherein the assay analysis includes a non-toxicity analysis or an efficacy analysis, which is configured to predict the non-toxicity or efficacy of one or more compounds applied to multiple feasible samples of cell structures in the in vitro microscopy assay, the method comprising: receiving a set of images associated with the multiple samples; inputting each image in the set of images into a first ML model, the first ML model being configured to predict the phenotypic features of the cell structure within the sample associated with each image; inputting each of the predicted phenotypic features associated with each sample into a second ML model, the second ML model being configured to predict a lower-dimensional phenotypic feature embedding for each sample; comparing the distance between the lower-dimensional phenotypic feature embedding of each sample and the lower-dimensional phenotypic feature embedding of a sample to which a compound with known non-toxicity or efficacy is applied; and based on the comparison, outputting, for each sample, an indication of the non-toxicity or efficacy of the compound applied to the sample.

[0020] As an option, the first ML model is a neural network or convolutional neural network (CNN) model. Optionally, the neural network or CNN model is trained using cell image training data for classification. As an option, the predicted phenotypic feature is embedded in a complete layer of the trained neural network or CNN model, and the method includes outputting the phenotypic feature from the complete layer. As an option, the final complete layer of the neural network or CNN model is used to output the embedding of the phenotypic feature.

[0021] A computer-implemented method as described in the first aspect, wherein the second ML model is based on a uniform manifold approximation and projection (UMAP) algorithm or a t-SNE algorithm for dimensionality reduction of phenotypic feature embeddings of the sample, wherein the phenotypic feature embeddings are mapped to a lower dimensional vector space for comparing the distance between the phenotypic feature embeddings and the phenotypic feature embeddings of samples with compounds of known toxicity, wherein the second ML model is trained on the UMAP technique using unsupervised training based on negative and positive control samples from the multiple samples to predict a toxicity distance metric associated with the sample to which the compound with known toxicity is applied.

[0022] A computer-implemented method as described in the first aspect, wherein the second ML model is based on a UMAP algorithm or a t-SNE algorithm for dimensionality reduction of phenotypic feature embeddings of the sample, wherein the phenotypic feature embeddings are mapped to a lower dimensional vector space for comparing the distance between the phenotypic feature embeddings and the phenotypic feature embeddings of samples having compounds with known non-toxicity or efficacy, respectively, wherein the second ML model is trained on the UMAP technology using unsupervised training based on negative and positive control samples from the multiple samples to predict a toxicity distance metric associated with samples to which compounds with known non-toxicity or efficacy are applied, respectively.

[0023] As an option, training the second ML model includes iteratively performing a grid search over the set of hyperparameters of the UMAP technique to select those hyperparameters that maximize the difference between the negative control samples and the positive control samples.

[0024] A computer-implemented method as described in the first aspect, wherein indicating the toxicity, non-toxicity or efficacy of the phenotypic feature embedding of the sample to which the compound is applied, respectively, comprises applying the phenotypic feature embedding of the sample to which the compound is applied to a second ML model to output a lower dimensional embedding of the sample to which the compound is applied; and determining an indication of the toxicity, non-toxicity or efficacy of the sample to which the compound is applied based on comparing the distance between the corresponding lower dimensional embedding and the embedding of one or more samples with known toxicity, non-toxicity or efficacy to which the compound is applied.

[0025] Optionally, indicating toxicity, non-toxicity or efficacy of the phenotypic feature embedding of the sample to which the compound is applied, respectively, further includes: applying the phenotypic feature embedding of the sample to which the compound is applied to a second ML model to output a lower dimensional embedding of the sample to which the compound is applied; and applying the lower dimensional embedding of the sample to which the compound is applied to a third ML model trained to output an indication of the distance between the lower dimensional embedding and a set of lower dimensional embeddings associated with negative control samples.

[0026] The computer-implemented method of the first aspect further comprises training a third ML model based on performing a grid search on a hyperparameter set of a high-dimensional distance metric algorithm that maximizes the distance between the lower dimensional embeddings of the negative control samples and the positive control samples while minimizing the distance between the lower dimensional embeddings of the negative control samples or minimizing the distance between the lower dimensional embeddings of the positive control samples.

[0027] In the computer-implemented method of the first aspect, the distance metric is the Wasserstein distance metric, and the high-dimensional distance metric algorithm is the Sinkhorn algorithm for estimating the Wasserstein distance between embeddings.

[0028] The computer-implemented method of the first aspect, wherein the distances for comparing the toxicity of the phenotypic feature embeddings indicating the sample are based on a Wasserstein distance metric. The computer-implemented method of the first aspect, wherein the distances for comparing the non-toxicity of the phenotypic feature embeddings indicating the sample are based on a Wasserstein distance metric. The computer-implemented method of the first aspect, wherein the distances for comparing the efficacy of the phenotypic feature embeddings indicating the sample are based on a Wasserstein distance metric.

[0029] According to a second aspect, there is provided an apparatus comprising a processor, a memory unit and a communication interface, wherein the processor is connected to the memory unit and the communication interface, wherein the processor and the memory are configured to implement the computer-implemented method of the first aspect.

[0030] According to a third aspect, there is provided a computer readable medium comprising data or instruction codes which, when executed on a processor, cause the processor to perform the computer implemented method of the first aspect.

[0031] According to a fourth aspect, a tangible computer-readable medium is provided, comprising data or instruction code for identifying viable samples of cell structures for analysis in in vitro microscopy, which, when executed on one or more processors, causes at least one of the one or more processors to perform at least one of the steps of the method: automatically identifying a first sample set that can be used for analysis from a plurality of samples in an assay plate; generating a 2-dimensional 2D image set for each sample in the first sample set, the 2D image set of each sample comprising a plurality of 2D image slices taken along the z-axis of each sample; identifying a viable sample set from the 2D image slice set; and outputting data representing the viable sample set for analysis.

[0032] According to a fifth aspect, a system is provided, comprising: a sampling module, which is configured to identify a first sample set that can be used for analysis from a plurality of samples in an assay plate; an imager module, which is configured to generate a 2-dimensional 2D image set for each sample in the first sample set, wherein the 2D image set of each sample includes a plurality of 2D image slices taken along the z-axis of each sample; a sample feasibility module, which is configured to identify a feasible sample set from the 2D image slice set; and an output module, which is configured to output data representing the feasible sample set for analysis.

[0033] In various embodiments, a computer program instruction optionally stored on a non-transitory computer-readable medium, which, when executed by one or more processors of a data processing device, causes the data processing device to execute these program instructions so that the one or more processors perform operations including one or more aspects of the embodiments described above and / or below (including one or more aspects of the appended claims).

[0034] In various embodiments, an apparatus is disclosed, the apparatus comprising a computer-readable storage medium and one or more processors, the computer-readable storage medium having program instructions contained therein, the one or more processors being configured to execute the program instructions to cause the apparatus to perform operations including one or more aspects of the embodiments described above and / or described below (including one or more aspects of the appended claims). The apparatus may include one or more processors or dedicated computing hardware. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to make the present invention more easily understood, embodiments of the present invention will now be described by way of example only with reference to the accompanying drawings, in which:

[0036] FIG. 1 a illustrates an example compound analysis pipeline system for performing downstream analysis from a quality controlled output sample of an in vitro microscopy assay according to some embodiments of the present invention;

[0037] FIG. 1 b illustrates an example of a quality control image analysis process for selecting and pre-processing microscopy images of cellular structures prior to use in training downstream efficacy, non-toxicity, or toxicity models and / or as input to trained downstream efficacy, non-toxicity, or toxicity models for predicting efficacy, non-toxicity, or toxicity, respectively, of a compound applied to a cellular structure in the sample of FIG. 1 a, according to some embodiments of the present invention;

[0038] FIG. 1 c shows an example of a first quality control process of a first quality control apparatus according to some embodiments of the present invention, the first quality control process being used together with the quality control image analysis process of FIG. 1 b to select a first feasible image set of a sample for analysis;

[0039] FIG. 1d illustrates an example of an image analyzer of a first quality control apparatus for use with the first quality control process of FIG. 1c according to some embodiments of the present invention;

[0040] FIG. 1e illustrates an example of a first quality control apparatus for use with the first quality control process of FIG. 1c according to some embodiments of the present invention;

[0041] FIG. 1f illustrates an example of a second quality control process of a second quality control apparatus according to some embodiments of the present invention, the second quality control process being used together with the quality control image analysis process of FIG. 1b to preprocess and select a final feasible image set of a sample for analysis;

[0042] FIG. 1g shows another example of a second quality control process controlled by a second quality control device according to some embodiments of the present invention, the second quality control process being used together with the quality control image analysis process of FIG. 1b for preprocessing and selecting a final feasible image set of a sample for analysis;

[0043] FIG. 2 a illustrates an example non-toxicity, efficacy or toxicity prediction system for predicting the non-toxicity, efficacy or toxicity, respectively, of a compound applied to the cellular structure of a sample for microscopy assay according to some embodiments of the present invention;

[0044] FIG2 b illustrates an example non-toxicity, efficacy or toxicity prediction process for use in the system of FIG2 a according to some embodiments of the present invention;

[0045] Figure 3 An example neural network for predicting phenotypic embeddings of cellular structures in the microscopy assayed sample of FIG. 2 a is shown according to some embodiments of the present invention;

[0046] FIG4 a illustrates an example assay plate having negative and positive control sample groups for training a deep learning toxicity model and a test sample group for input to the trained DL toxicity model according to some embodiments of the present invention;

[0047] FIG4 b illustrates an example unsupervised training process for a deep learning toxicity model using the negative and positive control sample sets of FIG4 a according to some embodiments of the present invention;

[0048] FIG. 4 c shows an example of using the trained deep learning toxicity model of FIG. 4 b to predict the toxicity of a compound according to some embodiments of the present invention;

[0049] FIG5 a illustrates another example assay plate having negative and positive control sample groups for training a deep learning toxicity model and a test sample group for input to the trained DL toxicity model according to some embodiments of the present invention;

[0050] FIG5 b shows an example distance matrix of negative and positive control samples for the trained DL model of FIG5 a according to some embodiments of the present invention;

[0051] FIG5 c shows another example distance matrix of negative and positive control samples and test samples for predicting the toxicity of a compound of a test sample using the trained DL toxicity model of FIG5 b according to some embodiments of the present invention;

[0052] Figure 5d shows an example of the conventional toxicity prediction method used on a set of 14 compounds applied to a sample of cell structures;

[0053] FIG5 e shows an example of conventional toxicity prediction results for a compound set from the conventional toxicity prediction method of FIG5 d ;

[0054] FIG5 f shows an example of the toxicity prediction results of the trained DL toxicity model for the same set of compounds as FIG5 e ;

[0055] Figure 6a is a schematic diagram of a system / apparatus for performing the methods described herein;

[0056] Figure 6b is a schematic diagram of another example system for performing the methods described herein; and

[0057] Figure 6c is a schematic diagram of another example system for performing the methods described herein.

[0058] Common reference numerals are used throughout the drawings to indicate similar features. DETAILED DESCRIPTION

[0059] Various example embodiments described herein relate to methods, devices and systems for automatically, efficiently and reliably performing quality control on images captured from samples of cell structures output from HTS microscopy assays. HTS microscopy assays output image sets of feasible samples of cell structures, wherein at least one set of samples assayed has been perturbed by one or more tested compounds. Compounds can be non-toxic or toxic compounds. The received feasible image sets can be applied to downstream analysis processes, such as but not limited to, for example: a deep learning (DL) model trained and configured to predict the non-toxic effects of compounds on perturbed cell structure samples; a DL model trained and configured to predict the efficacy effects of compounds on perturbed cell structure samples; a DL model trained and configured to predict the toxic effects of compounds on perturbed cell structure samples; or any other type of DL model trained and configured to predict the effects of compounds on perturbed cell structure samples.

[0060] The cell structure of the sample is associated with an organ of a subject or patient and can be designed to mimic or simulate the organ. The cell structure can be, but is not limited to, for example, based on at least one of the group consisting of: a cell spheroid; a vesicle; an organoid; a cell structure of an immortalized cell line; and any other suitable cell structure that mimics or simulates one or more processes of an organ of a subject or patient.

[0061] The DL model of the compound analysis pipeline system can be trained and applied to any type of cell structure associated with any organ and / or any related disease, any cell line, cell array and / or cell sample well to which a compound has been added, and used to predict at least one of the non-toxic effect, efficacy effect and / or toxic effect of the compound on the cell sample. For example, the compound analysis pipeline system can be applied to cell structure samples that mimic organs, such as but not limited to, for example, lungs, skin, kidneys, pancreas, liver, cardiac cell structures / heart, neural cell structures and / or any other organ of a subject or patient. Such cell structures can be used in samples to which compounds are applied in in vitro microscopy assays, etc. and automatically analyzed for non-toxicity, efficacy or toxicity by the compound analysis pipeline system.

[0062] FIG. 1a illustrates an exemplary compound analysis pipeline system 100 for automatically predicting the non-toxicity, efficacy or toxicity of one or more compounds applied to multiple samples of cell structures in an in vitro microscopy assay. The compound analysis pipeline system 100 includes an in vitro HTS microscopy assay system 102, a quality control imaging analysis system 104, and a compound prediction system 106, for example, predicting the non-toxicity, efficacy or toxicity of the compound. The in vitro HTS assay system 102 is configured to acquire a sample set of cell structures 102a and apply one or more compounds or reagents 102b to the sample set of cell structures 102a to input into a well sample set in a microscopy assay plate 102c for HTS staining and microscopy assay imaging 102d. In HTS staining and microscopy assay imaging 102d, the sample wells can be treated / stained with fluorescent reagents / compounds to emphasize the cell structure of each sample. For example, the treated / stained cell structures (such as, for example, spheroid structures, vesicles and / or cell nuclei) can be imaged by a microscopy imaging system. Thus, the in vitro HTS microscopy assay system 102 is configured to output a well image sample set 103. Each well sample image in the set corresponds to each sample of the cell structure and applied compound assayed in the well sample set.

[0063] The quality control imaging analysis system 104 can be configured to receive the well sample image set and perform image preprocessing and / or analysis to identify which samples from the well sample set are feasible for further downstream analysis, such as, for example, a) a prediction of non-toxicity of compounds applied to samples of the well sample set; b) a prediction of efficacy of compounds applied to samples of the well sample set; or c) a prediction of toxicity of compounds applied to samples of the well sample set; or d) any other type of property / compound prediction. In the quality control imaging analysis system 104, one or more image processing and / or machine learning algorithms can be applied to identify the feasibility of each well sample image based on any detected image artifacts and / or imaging defects, etc., and / or to enhance or emphasize the cell structure of interest within each feasible well sample image in the well sample set. Therefore, the imaging analysis system 104 can output a feasible well sample image set 105 for further downstream analysis. In essence, the feasible well sample image set 105 can be any suitable image set of cell structures from the assayed well sample set that fully describes the cell structure of the sample for automatic down analysis.

[0064] The downstream analysis system 106 is configured to receive a set of cell structure images 105 from a plurality of cell structure samples to which one or more compounds are applied during an in vitro microscopy assay. In this case, the received set of sample images 105 may be a set of viable well sample images 105 output from the imaging analysis system 104.

[0065] In the downstream analysis system 106, the received sample image sets 105 can each be input into a deep learning (DL) compound analysis and prediction model, which is configured to predict the toxicological, non-toxicological, or efficacy effect of each corresponding compound applied in the assay from the in vitro HTS microscopy assay system 102. The DL compound analysis and prediction model can be configured to predict the toxicity of each corresponding compound applied in the assay. The DL compound analysis and prediction model can be configured to predict the non-toxicity of each corresponding compound applied in the assay. The DL compound analysis and prediction model can be configured to predict the efficacy of each corresponding compound applied in the assay.

[0066] In any case, the DL compound analysis and prediction model can be based on any one or more DL modeling techniques / algorithms and / or machine learning (ML) techniques / algorithms that have been used to train the DL compound analysis and prediction model to identify or predict whether each of the received sample images 105 indicates the requested effect (e.g., toxicity, non-toxicity, or efficacy), even when the compound has not yet been applied to the cell structure of one or more of the well samples. The one or more DL / ML techniques / algorithms can be based on supervised ML, unsupervised ML, and / or semi-supervised ML algorithms, etc. However, for the task of training a DL compound analysis and prediction model related to toxicity, non-toxicity, or efficacy, it has been found that supervised learning is difficult due to the limited number of labeled training data sets related to cell structures that indicate toxicity or non-toxicity or efficacy, depending on whether the compound has been applied. Therefore, a combined supervised / unsupervised DL model training architecture can be used to train one or more component models of the DL compound analysis and prediction model.

[0067] For example, in the example of FIG. 1a, the DL compound analysis and prediction model may include a trained machine learning (ML) phenotypic feature extraction (FE) model 106a for extracting phenotypic features of a cell structure from each received sample image of the received sample image set 105. Supervised learning can be used to train the ML phenotypic FE model. The ML phenotypic FE model may be based on, but not limited to, for example, a neural network (NN) classifier that is trained using supervised training of readily available labeled / annotated training data sets to classify images of cells, organoids, spheroids, cell structures, etc. (e.g., images of cell structures are classified to determine whether the cells are cancer cells or tumor cells). The NN classifier may be based on any NN structure, such as, but not limited to, for example, a feedforward NN (FNN), a recursive NN (RNN), an artificial NN (ANN), a convolutional NN (CNN), any other type of NN, modifications thereof, or combinations thereof. Before the output classification layer or SoftMax output of the NN classifier, the phenotypic representation of the cell structure may be embedded through the high-dimensional output of one of the hidden layers or complete layers of the NN classifier. The NN classifier is configured to output an embedding of the phenotypic representation of the cell structure from the hidden layer or the complete layer, rather than outputting a classification.

[0068] Typically, the phenotypic representation embedding of the input image samples from the received image sample set 105 is a high-dimensional representation of the phenotypic features (e.g., for a CNN-type NN classifier / model, the dimensionality may be on the order of 2024 or greater). The trained NN classifier may be used to output a high-dimensional phenotypic feature representation of the cell structure of each input well sample to which the compound was applied in the sample image set 105. As an example, the neural network classifier may be based on a convolutional neural network (CNN) for classifying images of cell structures (e.g., classifying images of cell structures to determine whether a cell is a cancer cell), wherein, once trained, one of the final complete layers of the CNN may be used as a phenotypic feature representation of each image sample in the received image sample set 105 input thereto.

[0069] The ML phenotype embedding model is coupled to a trained ML low-dimensional (LD) embedding model that is trained and configured to optimally embed a high-dimensional phenotypic representation of a predicted phenotypic feature of each received sample image of the received image sample set 105 into a lower-dimensional phenotypic embedding in a lower-dimensional space for further analysis, wherein the lower-dimensional phenotypic feature embedding includes those phenotypic features associated with toxicity, non-toxicity, or efficacy according to the analysis type. Rather than using supervised training on the ML LD embedding model, one or more unsupervised DL / ML techniques / algorithms are used to ensure that the trained ML LED embedding model outputs an LD phenotypic embedding that represents phenotypic features associated with toxicity, non-toxicity, or efficacy of the compound to the cell structure according to the analysis type. The unsupervised DL / ML techniques / algorithms can be based on at least a clustering algorithm or a dimensionality reduction algorithm, such as, but not limited to, for example, a support vector machine (SVM), a uniform manifold approximation and projection (UMAP), or a t-distributed stochastic neighbor embedding (t-SNE) type algorithm, a combination thereof, a modification thereof, and the like.

[0070] For example, the ML LD embedding model can be trained using negative control samples and positive control samples included on the assay plate 102c in an in vitro microscopy assay. That is, the assay plate 102c can include a negative control sample group and a positive control sample group, the negative control sample group can be a first set of wells of the assay plate 102c, wherein the sample is not perturbed or has no compound applied thereto, and the positive control sample group can be a second set of wells of the assay plate 102c, wherein the sample is perturbed by a known compound with known toxicity, known non-toxicity, or known efficacy according to the type of analysis. The assay plate 102c can also include a third set of wells of the assay plate 102c, the third set of wells including a cell structure sample with a compound applied thereto. Therefore, the negative image group and the control image group associated with the negative control sample and the positive control sample can be used to train the ML LD embedding model in an unsupervised manner.

[0071] For example, a negative image group and a control image group are input to an ML phenotype embedding model, which outputs a corresponding negative high-dimensional phenotype representation and a positive high-dimensional phenotype representation. Then, an iterative optimization of the parameters of a UMAP or t-SNE algorithm can be performed on the high-dimensional negative and positive control phenotype representations output from the ML phenotype embedding model 106a, wherein the difference between the resulting LD negative and positive control phenotype representation embeddings is maximized. The resulting optimized parameters can be used with a UMAP or t-SNE algorithm for dimensionality reduction of a high-dimensional phenotype representation corresponding to a third set of well samples. The LD phenotype representation can be output for comparison with a negative control (NC) LD phenotype representation using a suitable distance or similarity metric used by a trained ML distance model.

[0072] The ML LD embedding model is coupled to a trained ML distance model that is configured to perform a distance or similarity metric comparison between the LD phenotype representation of one of the image samples of the well and the negative control LD phenotype representation using a suitable distance or similarity metric associated with the LD space of the LD phenotype representations. It should be noted that the LD space into which the LD phenotype representations output by the ML LD embedding model are embedded can still be considered a high-dimensional space that cannot be reliably used by the Euclidean distance metric or similarity metric decomposition. Therefore, a high-dimensional distance metric or similarity metric can be used based on, but not limited to, for example, the Wasserstein distance and / or any other high-dimensional distance metric or similarity metric. The ML distance model outputs a prediction of toxicity, non-toxicity, or efficacy as a probability based on a distance comparison (e.g., a Wasserstein distance comparison) between each LD phenotype representation of the sample and the negative control LD phenotype representation.

[0073] FIG. 1 b illustrates an example of a quality control process 110 in a quality control image analysis and feature extraction system 104 for the compound analysis pipeline system 100 of FIG. 1 a, which is used to automatically identify feasible analyzable images of images of sample 103 captured from the HTS in vitro microscopy system 102 of FIG. 1 a. The QC process 110 is used to determine which samples or groups of samples from the set of wells on the assay plate are analyzable and non-analyzable. The QC process 110 uses ML techniques to help determine which wells' image captures can be retained or discarded for downstream analysis. The feasible image set is analyzable because the cell structure is sufficiently well defined for downstream image analysis processing, such as, but not limited to, for example, extracting phenotypic feature representations from the cell structure of the image for training and / or input to a DL / ML model of a downstream analysis system 106, as described with reference to FIG. 1 a to 5 e. For example, the feasible image set 105 output from the QC process 110 improves the robustness and reliability of the trained DL model of the downstream analysis system 106, which further enhances the robustness and accuracy of the toxicity, non-toxicity, or efficacy prediction output by each of the toxicity, non-toxicity, or efficacy prediction models of the downstream analysis system 106 relative to the toxicity, non-toxicity, or efficacy prediction of the test sample to which the compound is applied.

[0074] Although the set of possible images 105 output by the QC process 110 is described with reference to the downstream analysis system 106, this is by way of example only, and the skilled person will appreciate that the set of possible images 105 output by the QC process 110 may be input into any other downstream process associated with analyzing, classifying, and / or predicting / estimating one or more aspects or characteristics of the images, depending on the type of HTS in vitro microscopy assay being performed. For example, an HTS in vitro microscopy assay may be used in a drug / compound research program to determine a characteristic or effect of one or more compounds on a cellular structure sample for an in vitro microscopy assay, such as, but not limited to, toxicity of the compound when applied to a cellular structure sample, non-toxicity of the compound when applied to a cellular structure sample, efficacy effect of the compound on a cellular structure sample, and / or any other type of characteristic of a compound when applied to a cellular structure sample in an in vitro microscopy assay, etc.

[0075] In this example, the QC process 110 identifies a first feasible image set of samples from images 103 captured by the HTS in vitro microscopy system 102 and preprocesses the first feasible image set into a final feasible image set 105 for input to (but not limited to), for example, a downstream analysis system 106 and / or any other compound analysis system or downstream workflow process / analysis system. With reference to the downstream analysis system 106, the QC process 110 outputs the feasible image set 105 of the samples for training the DL model of the downstream analysis system 106 of FIG. 1a and / or as input to the trained DL model of the downstream analysis system 106, as described with reference to FIGS. 1a to 5e. The downstream analysis system 106 receives the feasible image set associated with the plurality of samples, wherein at least one group of the plurality of samples has a compound applied thereto for testing and / or as a positive control group for training the DL model of the downstream analysis system 106.

[0076] Upon receiving a set of images captured from a plurality of samples from an assay plate (e.g., assay plate 102c, 400, or 500 of FIG. 1a, 4a, or 5a), the QC process 110 identifies viable samples of cell structures from an in vitro microscopy assay for further downstream analysis. The QC process 110 includes the following steps:

[0077] In step 111, a first set of sample images that can be used for analysis is automatically identified from the received set of sample images captured from the plurality of samples of the assay plate. For example, the received set of sample images is automatically analyzed to determine whether features of the spheroid / cell structure of the sample are present and whether these features highlight possible artifacts and / or other imaging defects on the image or outside the focus area. Artifacts can bias predictions relative to downstream analysis using (multiple) ML models, etc. For example, the downstream analysis system 106 uses the received image of the sample 105 to predict whether the compound is toxic or non-toxic, so any artifacts may be detrimental to the training and / or prediction of toxicity or non-toxicity because they may mask the effect of the compound applied to the cell structure. For example, these artifacts can mask the dose at which the compound is actually toxic or non-toxic (e.g., a dose with a 50% toxic effect).

[0078] Automatic identification uses one or more ML models to estimate and predict whether the sample of each hole is analyzable. For example, automatically identifying a first sample image set can include inputting a received sample image set from an assay into one or more ML models, which are configured to identify one or more regions of each image in the set that may have a cell structure therein. The image set with the identified regions can be input into one or more additional ML models to identify whether the cell structure within the identified region of each image is suitable for further downstream analysis. Those images from the received image set having regions containing cell structures determined to be analyzable are selected to form a first sample image set that is feasible for further downstream analysis.

[0079] In step 112, for each image of a sample in the first sample image set, a 2-dimensional (2D) image set of the sample is generated, wherein the 2D image set of the sample includes a plurality of 2D image slices capturing the sample in an assay plate, wherein the plurality of 2D image slices are taken along the z-axis of the sample. The 2D image set of each sample in the first image set of samples is taken at a different z-axis position so as to form a 3D representation of each sample. The plurality of 2D images of each sample form a 3D representation of each well of each sample, which 3D representation can be further processed and represented as a fused / compressed 2D representation.

[0080] In step 113, a feasible sample set is identified from the 2D image slice set. For example, for each sample in the first sample image set, the 2D image slice set can be fused, compressed or combined to form a 2D representation of the 3D structure of the sample, wherein the 2D representation enhances the cell structure within the sample. The cell structure of each fused 2D representation can be further analyzed to identify whether they are sufficiently distinguishable or well defined for downstream analysis (e.g., for input to the toxicity prediction system 106). For example, when it is determined that the quality of the image is suitable for extracting one or more metrics or phenotypic representations of the cell structure from the image, the fused 2D representation is sufficiently distinguishable or well defined. Those samples from the first sample image set with the 2D image slice set that are identified as sufficiently distinguishable or suitable for further analysis form the feasible sample set for analysis.

[0081] In step 114, data representing the feasible sample set for analysis is output as an image set 105. For example, each fused 2D representation image of the feasible sample set can form the image set 105 of FIG. 1a, which can be input to a downstream process for analysis, such as a toxicity prediction system 106.

[0082] FIG. 1c illustrates an example of a first quality control (QC) process 115 for step 111 of the QC process 110 of FIG. 1b for selecting a first feasible image set of a sample suitable for further downstream analysis. The first QC process 115 is configured to automatically identify a first sample image set that is feasible for analysis, for example, for input to a downstream analysis system 106, as described with reference to FIGS. 1a to 5e. The first QC process 115 may include the following steps: In step 116, each image of a sample of the received sample image set 103 from the HTS in vitro assay system 102 is preprocessed to form a preprocessed image set of the sample. Preprocessing may include performing image analysis and / or processing based on, but not limited to, for example, correcting illumination, lighting, illumination field correction, artifact reduction / interpolation, or any other image processing algorithm for processing an image to further enhance a cellular structure that may be contained within the image. In step 117, the set of preprocessed images of the sample is input to a machine learning (ML) region of interest (ROI) model, which is trained and configured to identify one or more regions of interest in the input image of the sample for each input image of the sample, wherein the region of interest includes a cellular structure of the corresponding sample. The ML region of interest (ROI) model can be based on any type of neural network (NN) structure, such as but not limited to, for example, a feedforward NN (FNN), a recursive NN (RNN), an artificial NN (ANN), a convolutional NN (CNN), any other type of NN suitable for identifying a region of interest associated with a cellular structure in an image, modifications thereof, or combinations thereof. For example, the MLROI model can be based on a CNN that is trained to identify a ROI associated with a cellular structure within an input image and / or an output image focused on an identified ROI associated with a cellular structure within an input image. The MLROI model can be trained using a CNN structure and a labeled or annotated training image dataset, wherein each image in the training image dataset is labeled or annotated with information associated with whether a cellular structure is present and, if so, a region of interest within the image.

[0083] At this stage, those images in the input sample image set where the ML ROI model does not detect a region of interest may be discarded from the sample image set, as these images are more likely not to be analyzed. Alternatively, for each image where the ML ROI model does not detect a region of interest, the region of interest may be defaulted to the entire input image for further analysis in step 118. That is, the region of interest may be identified as the entire input image, which may be input to step 118.

[0084] In step 118, each of the images having the identified region of interest containing the cell structure of the sample is input to an ML image feasibility model that is trained and configured to classify whether the cell structure in the input image is analyzable. Thus, each input image can be classified with a label or probability value representing the feasibility of the analyzable input image. If the label or probability value indicates that the input image is analyzable (e.g., the probability value may be greater than or equal to a predetermined feasibility probability threshold, or the label indicates that the image is feasible), the input image is placed in the first set of sample images that are considered feasible. Remaining input images with labels indicating that the image is not feasible or with probability values ​​less than a predetermined feasibility probability threshold are discarded.

[0085] The image feasibility model can be based on, but not limited to, any ML algorithm or technique that can be used to classify, for example, whether the cell structure within the ROI in the image is analyzable. This can be based on identifying from each input image whether there are any phenotypic features associated with the cell structure (e.g., spheroids, nuclei, vesicles, etc.) within the input image, and evaluating whether there are sufficient phenotypic features in the cell structure of the sample captured in the image will enable further analysis of the cell structure of the sample in the image. For example, the ML algorithm used to train the image feasibility model can be based on, but not limited to, for example, a support vector machine (SVM) or a NN classifier / structure, etc. As an example, the image feasibility model can be based on a set of SVMs, each SVM being associated with classifying a specific phenotypic feature of the cell structure of the sample captured in the input image or a cluster of phenotypic features, wherein each SVM output indicates whether there is a positive or negative classification of a specific phenotypic feature. If enough SVM outputs indicate that there is a positive classification of the corresponding specific phenotypic feature, the input image can be considered to be an image feasible for analysis.

[0086] In step 119, a first sample image set including data representing those sample images classified as analyzable in step 118 is output. The output first sample image set may be further processed, such as but not limited to, for example, in steps 112-114 of the QC process 110 to further enhance the cellular structure contained in each image and / or for further selection related to the feasibility of the first sample image set. Alternatively or as an option, the first sample image set output from the first QC process 115 that is deemed feasible may be output from the image analysis system 104 as image set 105 for input to a toxicity prediction system 106 or other downstream process for analysis, etc.

[0087] FIG1d illustrates an example of an image analyzer 120 for implementing steps 116 and 117 of the first QC process 115 of FIG1c, wherein each image of the sample is pre-processed and a region of interest containing cellular structures within each image of the sample is identified. In this example, each image in the sample image set 103 output from the HTS in vitro microscopy assay system 102 is pre-processed by an image pre-processing unit 122. In this example, each input image 122a to the image pre-processing unit 122 has an illumination function 122b applied thereto, and the resulting illumination corrected image 122c output from the first image pre-processing unit 122 is passed to an image ROI unit 124. The illumination function 122b can be configured to correct for uneven fluorescence illumination of an image capture device in the HTS in vitro microscopy system 102. Although the image preprocessing unit 122 performs illumination field correction on each image, this is only an example and the present invention is not limited to this. Those skilled in the art should understand that the image preprocessing unit 122 can perform one or more other image processing functions, such as image focusing, illumination field correction, sharpening, saturation, artifact reduction, correction and / or interpolation, any other image processing algorithm, combination thereof, modification thereof, etc., wherein the cell structure of the sample captured for each input image is further enhanced for use by the image ROI unit 124.

[0088] Each image in the set of preprocessed images of the sample output from the image preprocessing unit 122 is input to the image ROI unit 124. The image ROI unit 124 is configured to include an ML region of interest (ROI) model based on a CNN architecture (e.g., CNN (CellPose (RTM)) based segmentation), which is trained and configured to identify each segment of the image and determine which segments contain cell structures, wherein the segments containing cell structures form the ROI of the image. The CNN-based segmentation (e.g., CellPose) architecture can be used on microscopy images to segment cell bodies, membranes, and nuclei. For example, CellPose is a deep learning CNN architecture that has been trained on a highly varied cell image dataset containing more than 70,000 segmented objects. These can be retrained and / or fine-tuned on an image training dataset focused on cell structures of microscopy assay images intended for downstream analysis. Although a CNN-based Cell Pose segmentation architecture is described, this is only an example and the present invention is not limited thereto, and the skilled person will appreciate that any other suitable ML algorithm / model may be used that performs identification of a ROI containing a cell structure associated with an input image of a sample. The ML ROI model may be trained using a labeled training dataset comprising a plurality of images using a CNN architecture, each of which is annotated with a label or annotation that includes data indicating whether a cell region of interest exists and / or the location of the region of interest within the image, etc. For example, a training image dataset may include images labeled and annotated with information associated with the presence or absence of a cell structure (e.g., phenotypic features of a cell structure and / or a macroscopic cell structure, cell spheroids, cell nuclei, cell mutations, perturbations, cell cholestasis, related biological features, etc.), and if a cell structure, a perturbed cell structure, or a remaining portion of a cell structure exists, the region of interest contains a cell structure within the image. Therefore, the ML ROI model is trained and configured to identify / predict a fragment / ROI of a remaining / mutated / perturbed cell structure in an image of an input image containing a sample and output the information for downstream processing.

[0089] For example, an input image of sample 124a may contain cell structures, such as cell spheroid objects containing cell nuclei, wherein only those segments of the image containing the cell nuclei are identified and form regions of interest of the image. In another example, an input image of sample 124b may contain cell spheroid objects representing cell structures undergoing cholestasis, i.e., the cell structures have been perturbed / mutated by a compound, and, but not limited to, for example, the cell nuclei may disappear, etc. In such a case, the ML ROI model has been trained to identify segments / ROIs of the image containing the remaining / mutated / perturbed cell structures in the image of the sample. The image ROI unit 124 may be configured to output only those portions of the image 124a containing the ROI. Alternatively or in addition, the image ROI unit 124 may be configured to input an image 124a annotated with data representing the ROI for use by other processing units / downstream processing when focusing on the cell structures contained in the ROI of the image. In any case, the image ROI unit 124 outputs an image data set representing each pre-processed input image set of the sample, and for each image in the set, the ROI contains the cell structure of the sample within each input image. The output image and ROI data of the image analyzer 120 can be used in steps 118 and 119 of the first QC process 115, which can be implemented as an image feasibility system configured to extract features in the ROI of each image of the sample and determine whether each image is of good quality or poor quality, that is, whether it is feasible or not feasible for further analysis, or whether each image is analyzable or not analyzable. The output of the image feasibility system will include an image set of analyzable / feasible samples and can form the image set 105, which can be input to further downstream processes, such as, but not limited to, for example, a toxicity prediction system 106.

[0090] FIG. 1e illustrates an example of an image feasibility system 130 as part of a QC pipeline for implementing steps 118 and 119 of the first QC process 115 of FIG. 1c for classifying whether the cell structure in the ROI of each input image of the sample is analyzable (e.g., feasible for further downstream analysis / processing). The image feasibility system 130 may include an image analyzer 120 of FIG. 1d that receives the microscopy image set of the sample 103 output from the HTS microscopy measurement system 102 of FIG. 1a and, for each image in the image set 103, outputs data representing each image and an ROI containing cell structure features of each image. The image set and corresponding ROI are input to a phenotype sampler 132 that identifies and / or classifies, for each image in the image set, whether the ROI of each image has n phenotypic features (e.g., P1, P2, P3, ..., Pn). Phenotypic clustering 134 is performed on each of the n phenotypic features detected within the ROI of each image of the sample. Each of the n phenotypic features has a class of SVMs 136a-136n trained to classify the ROI of the image relative to one of the n phenotypic feature clusters. Each of the SVMs 136a-136n outputs a classification of whether the ROI of the image has a corresponding phenotypic feature P1, P2, ... Pn, respectively. Each cluster of n phenotypic features P1, P2, P3, ..., Pn is input to a corresponding SVM 136a-136n, and each SVM outputs a probability or possibility that the corresponding phenotypic feature is located within the ROI of the image of the sample. These n output probabilities are combined to provide an output classification value 138 indicating whether the image is analyzable (e.g., whether it is feasible). Those images whose classification values ​​138 indicate that the image is analyzable or feasible (e.g., the classification value 138 is greater than a feasibility threshold) are selected to form a feasible first image set of the sample.

[0091] FIG. 1f illustrates an example of a second quality control (QC) process 140 in steps 113-114 of a quality control (QC) process 110 that enhances and selects a final feasible image set of a sample for downstream analysis, such as, but not limited to, the input image set 105 of the downstream analysis system 106 of FIG. 1a. Once a feasible first image set of a sample has been output from step 111 of the QC process 110 of FIG. 1b or from the image feasibility system 130 of FIG. 1e, step 112 of the QC process 110 may be performed, wherein, for each image of the sample from the first feasible image set, a 3D representation of each well sample corresponding to the each image of the sample is captured by a microscopy imaging device. The 3D representation of each image of the sample is generated by a microscopy imaging device (e.g., ImageXpress (RTM)) that captures a 2-dimensional (2D) image set of the sample, wherein the 2D image set of the sample includes capturing a plurality of 2D image slices of the sample in the well samples in the assay plate. The multiple 2D image slices are taken at different z-foci or different foci along the longitudinal z-axis of the sample well. The 2D image set of each image of the sample in the first feasible image set is taken at different z-axis positions so as to form a 3D representation of each sample. The 2D image set of each sample forms a 3D representation of the image of each well corresponding to the sample from the feasible image set. Each 2D image set of each sample is enhanced by fusing or compressing / combining the 2D image set into a single 2D representation of the image of the sample to emphasize the cellular structure in each 2D image. The second QC process 140 for enhancing and selecting the final feasible image set of the sample for downstream analysis includes the following steps:

[0092] In step 142, for each 2D image set of the sample from the feasible image set of the sample, the foreground, background and multiple uncertain feature areas or categories of the cell structure in each 2D image slice in the 2D image slices from the 2D image set are identified. The multiple uncertain feature areas or categories include multiple uncertain foreground feature areas or categories and multiple uncertain background feature areas or categories. For example, for each image in the 2D image set, each x / y pixel value is divided into multiple categories, including but not limited to, for example, a foreground category, a background category and multiple uncertain foreground categories and multiple uncertain background categories.

[0093] In step 144, for each 2D image set of the sample, the identified foreground, background and multiple uncertain feature regions or categories of the 2D image slices in the 2D image set of the sample are iteratively combined to generate a single 2D image of the cellular structure of the sample. For example, the iterative process takes each foreground category and mixes it with the uncertain foreground category to try and optimize the projection between the two categories in the iterative process by smoothing the projection in the local area. The iterative optimization process can have two criteria, the final projection needs to be locally smooth in terms of z-projection in the local area around the foreground category, and the local intensity in the local area is the maximum possible intensity. The iterative combination can be based on an ML smoothing model that iteratively smoothes the foreground category and the uncertain foreground category.

[0094] In step 146, for each single 2D image of the sample, the iteratively generated single 2D image of the sample is selected for the final feasible image sample set based on the quality of the single 2D image of the sample. For example, the quality of the single 2D image sample can be evaluated based on whether there are any sharp deviations between adjacent foreground feature regions and adjacent uncertain foreground regions, or whether there are any sharp deviations between adjacent background feature regions and adjacent uncertain background regions, etc. For example, if there are too many "jumps" between adjacent foreground regions / categories and adjacent uncertain foreground regions / categories at the end of the iterative process, or the resulting single 2D image still does not meet the criteria related to the image intensity of the local area of ​​the z projection, or if the iterative process does not converge to within an appropriate error threshold, the single 2D image is not feasible and can be selected for discarding. The remaining single 2D images of each sample can be used to form the final feasible image set.

[0095] In step 148, a final feasible image set is output that includes image data representing a single 2D image for each sample selected in step 146. The final feasible image set may form the image set 105 that is input to a downstream process, such as but not limited to, for example.

[0096] FIG. 1g illustrates another example of a smooth manifold extraction (SME) system 150 for implementing steps 112-114 of the QC process 110 of FIG. 1b and / or steps 142-148 of the second QC process 140 of FIG. 1f for further enhancing and selecting a final set of viable images 167a-167p for output from the quality control image analyzer system 104. The SME system 150 includes a chain of processing units including a 3D image representation unit 152, a parsing unit 154, an FFT unit 156, pixel classification units 158-160 for foreground / background classification / clustering, an iterative SME unit 162, and a final image selection unit 164 and an output unit 166. The output unit 166 outputs a final set of viable images 167a-167p that can be used as an input to the image set 105 for further downstream analysis processes, such as, but not limited to, for example, to a downstream analysis system 106.

[0097] The 3D image representation unit 152 is configured to perform step 112 of FIG. 1b, which receives a first feasible image set of the sample that has been output from step 111 of the QC process 110 of FIG. 1b or from the image feasibility system 130 of FIG. 1d. The 3D image representation unit 152 captures a 3D representation of the sample corresponding to each feasible image of the sample in each sample well 151. For example, a microscopy imaging device 153 (e.g., ImageXpress (RTM)) is directed to capture a 2-dimensional (2D) image set 155 of the sample well 151 corresponding to the feasible image of the sample. The 2D image set 155 of the sample includes a plurality of 2D image slices 155a-155k of the sample from the corresponding sample well 151 of the assay plate. The sample well 151 has a longitudinal z-axis 151a, an x-axis 151b, and a y-axis 151c, wherein the x-axis 151b and the y-axis 151c are orthogonal to each other and also to the longitudinal z-axis. The 2D image slices 155a-155k are taken at different z-foci or different focal points along the longitudinal z-axis 151a of the sample well 151. The 2D image set 155 of each possible image of the sample is taken at a different z-axis position so that a 3D representation of each sample is formed. The 2D image set 155 of each sample forms a 3D representation of each well 151 corresponding to a possible image of the sample from the received possible image set.

[0098] The remaining processing units 154-166 of the SME system 150 are configured to process each 2D image set 155 of each sample to enhance / emphasize the cellular structure of the sample represented in each 2D image. This is performed by fusing or compressing / combining each 2D image set 155 of each sample into a single 2D representation 165 of the sample. In essence, these remaining processing units 154-166 are configured to find, for each sample, the z-focus of each pixel of each 2D image set 155 corresponding to the relevant structure of the cell / cellular structure of each pixel.

[0099] In the analysis unit 154, a profile of each 2D image slice 155a in the 2D image set 155 is extracted, wherein any (x, y) position corresponds to a profile of a focus value in the z direction 151 of the sample hole 151, the profile being composed of a direct intensity value in the case of a confocal image or an SML value in a wide-field epi-fluorescence image. Each of the 2D image slices 155a-155k (e.g., profiles) that pass through a certain foreground signal contains a lower frequency component. The profile of each of the 2D image slices 155a-155k is passed to the FFT unit 156. For each 2D image set of the sample, the FFT unit 156 performs a fast Fourier transform (FFT) on each of the profiles of the 2D image slices within the 2D image set 155 of the sample to determine a power spectrum of each of the profiles of the 2D image slices 155a-155k. The FFT profiles of the 2D image slices 155a-155k are passed to the pixel classification units 158-160 for foreground / background classification and clustering.

[0100] The pixel classification units 158-160 use an ML classification algorithm / model to perform multi-class classification on the FFT profiles of the 2D image slices 155a-155k, thereby classifying each (x, y) location or pixel of each profile of the 2D image slices into one of a plurality of labels associated with a foreground class 159a, a background class 159b, and / or a plurality of uncertain foreground / background classes 159c-159l. The ML classification algorithm may be based on, but not limited to, for example, a multi-class k-means algorithm, a multi-class UMAP algorithm, and / or any other suitable multi-class classification algorithm or model capable of classifying the (x, y) location of each pixel of each FFT profile of the 2D image slices 155a-155k of the 2D image set 155 into a label associated with a foreground class 159a, a background class 159b, or a plurality of uncertain foreground / background classes 159c-159l.

[0101] The pixel classification unit 156 outputs a labeled / annotated FFT profile 161 representing 2D image slices 155a-155k to the SME unit 162, wherein each (x,y) pixel of each 2D image slice has been labeled with a classification associated with a foreground class 159a, a background class 159b, or one of a plurality of uncertain foreground / background classes 159c-159l.

[0102] The SME unit 162 is based on an ML SME model / algorithm 163 (or an iterative SME optimization process) that iteratively processes the data representing the 2D image set of each sample based on the determined foreground, background, and uncertain foreground and background categories of the pixels of the 2D image slice of the 2D image set assigned to each sample. The ML SME model 163 uses an iterative SME algorithm. The ML SME model 163 is configured to perform an iterative SME optimization process that combines the data representing the 2D image slice of the sample using a cost function that minimizes the balance between local smoothness and proximity to the maximum focus value, thereby obtaining a final smoothed index map 163m for the sample. It should be noted that the first index map 163a is highly discontinuous at the beginning of the iterative SME optimization process, but should be smoothed when the iterations of the SME optimization process converge to the final smoothed index map 163m while retaining fine details on the foreground. The final smoothed index map 165m is sent to the extraction and selection unit 164.

[0103] For example, the SME algorithm 163 of the SME unit 162 looks at foreground pixels mixed with uncertain foreground classes and attempts to optimize the projections of these types of foreground classes of the 2D image slices 155a-155k and fuse / compress them into a final single 2D image 165a (or final index map 163m) using an iterative SME optimization process. When a final error threshold has been reached / satisfied and / or a maximum number of iterations has been met / reached, the iterative SME optimization process smoothes the projection in each iteration from the first iteration (e.g., iteration 001) to the final iteration (e.g., iteration 171). The iterative SME optimization process needs to meet two criteria, the first criterion being that the final projection (e.g., final index map 163m or final single 2D image) needs to be locally smooth in terms of z-projection, and the second criterion being that the local intensity is the maximum possible intensity (e.g., using a Gaussian energy function). These criteria are used in each iteration from the first iteration to the final iteration when the 2D image slices are fused / compressed together to form the final single 2D image. Once the iterative SME optimization process of the SME algorithm 163 has converged to the final index map 163m and / or the maximum number of iterations has been met, the resulting final index map 163m is used to form the final single 2D image 165a. The resulting final index map 163m or the final single 2D image 165a is passed to the extraction and selection unit 164 to determine whether the final single 2D image 165a should be included in the final feasible image set.

[0104] In the extraction and selection unit 164, the final smoothed index map 163m output from the iterative SME algorithm 163 is analyzed, and a plurality of voxels corresponding to the index map 163m are extracted from the original stack to produce a final single 2D image 165a. The extraction and selection unit 164 using the selection unit 165b selects whether to include the final single 2D image 165a in the final feasible image set of the sample to be sent to the output unit 166. The selection unit 165b determines whether the final single 2D image 165a or the index map 163m corresponding to the final single 2D image 165a has sufficient quality or sufficient smoothness for output from the extraction and selection unit 164. For example, the selection unit 165b determines whether the iterative SME algorithm 163 or the final single 2D image 165a meets two criteria, the first criterion is based on determining whether the projection of the final single 2D image 165a or the final index map 163m of the iterative SME algorithm 163 is locally smooth in terms of z-projection, and the second criterion is based on determining whether the local intensity is at the maximum possible intensity.

[0105] For example, the first criterion is based on checking the smoothness of each local pixel area of ​​the final single 2D image 165a (or the final index map 163m). In this case, each pixel local area of ​​the final single 2D image 165a (or the final index map 163m), such as but not limited to, for example, each 5×5 pixel neighborhood in the final 2D image 165a (or the final index map 163m), the z-level in each local area (e.g., 5×5 pixel neighborhood) should be smooth in terms of z-level, where there should not be any deviation or "jump" in intensity when moving from 1 pixel to another pixel in the local area (e.g., not jumping directly from intensity level 1 to 17). The second criterion is based on the intensity of the z-projection of each local pixel area of ​​the final single 2D image 165a (or the final index map 163m). The second criterion is met when, within each local area (e.g., 5×5 pixel neighborhood), the z-projection obtained on the local area (e.g., 5×5 pixel neighborhood) has an intensity value based on the local area (e.g., 5×5 pixel neighborhood) at the maximum z-intensity. In the selection unit 165b, if these two main criteria are not met, the selection unit 165b discards the final single 2D image 165a (or the final index map 163m) that does not form part of the final feasible image set 167. Otherwise, the final single 2D image 165a (or the final index map 163m) of the sample is output by the extraction and selection unit 164 to be sent to the output unit 166, and the output final single 2D image 165a is added as the final feasible image 167a of the final feasible image set 167.

[0106] Although the above first and second criteria are described herein, this is only an example, and those skilled in the art will appreciate that modifications to these two criteria and / or other criteria may be applied to select / discard the final single 2D image 165a when it is sent to the output unit 166. For example, these modifications and / or other criteria may be based on, but not limited to, for example, the first criterion may be further modified to count the number of "jumps" of the intensity deviation for determining whether to discard / select the final single 2D image 165a (e.g., if the number of "jumps" of the intensity deviation is below a minimum "jump" threshold / count, the final single 2D image 165a may be selected for output, otherwise it may be discarded), and / or the selection unit 165b may further consider whether the iterative SME algorithm converges within a predetermined error threshold (e.g., if the convergence error of the iterative SME algorithm is less than or equal to the predetermined error threshold, the final single 2D image 165a may be selected for output, otherwise it may be discarded), and / or any other criteria for evaluating whether the final single 2D image 165a of each sample is of sufficient quality and / or can be analyzed by further downstream processes, etc.

[0107] The output unit 166 stores each of the output final single 2D images 165a selected for output by the extraction / selection unit 164 as a plurality of images forming a final feasible image set of the sample. The output unit 166 can then send the final feasible image set of the sample to a further downstream process, such as but not limited to, for example, as an input image set 105 of a downstream analysis system 106.

[0108] Although the output data representing the output final single 2D image 165a can be used to form the final feasible image set of samples, this is only an example and the present invention is not limited thereto, and the skilled person should understand that the feasible samples associated with each output final single 2D image in the output final single 2D image 165a can be identified, and therefore, any other image captured by each feasible sample can be used as the output feasible sample set, etc. For example, outputting data representing the image of the feasible sample set can further include outputting data representing one or more of the group consisting of: a 2D image set generated for each feasible sample in the feasible sample set; a pre-processed image of each feasible sample in the feasible sample set; a single 2D image generated for each feasible sample, each single 2D image generated by iteratively combining the 2D image slice of the feasible sample based on the identified foreground, background and uncertainty region of the 2D image slice of the feasible sample; and any other image captured or processed in relation to the feasible sample; a combination thereof, a modification thereof and / or a modification according to application requirements.

[0109] In the following description, the downstream analysis system 104 will be described with reference to various example embodiments of a toxicity prediction system, which relates to (multiple) methods, devices and (multiple) systems for automatically, efficiently and reliably testing and predicting the toxicity of compounds applied to cell structure samples in HTS microscopy assays. Although a toxicity prediction system is described, this is only an example and the present invention is not limited thereto, and the skilled person should understand that the concepts, DL / ML models and structures used in the toxicity prediction system can be applied to a non-toxic prediction system or an efficacy prediction system. These systems can be configured to receive a sample image set of a cell structure, for example, such as an image 105 output from a quality control image analysis system 104 as described with reference to Figures 1a to 1g, wherein at least one set of samples being assayed has been perturbed by one or more test compounds.

[0110] In the toxicity prediction system, the received image set is applied to a DL model that is trained and configured to predict the toxic effects of compounds on cell structure samples that have been perturbed. The DL model generates a high-dimensional phenotypic representation of the cell structure represented in each received image sample at the input. These high-dimensional phenotypic representations are further processed into lower-dimensional phenotypic embeddings, focusing on the features of cell structures associated with toxicity. A toxicity prediction for each image sample is output based on the distance between each lower-dimensional phenotypic embedding and the negative control lower-dimensional phenotypic embedding or a similarity measure used to estimate the distance, and the negative control is associated with a cell structure sample that has not been perturbed by any tested compound.

[0111] The toxicity prediction system is described herein with reference to drug-induced hepatotoxicity or drug-induced liver injury (DILI) by way of example only, but not limitation, as an acute or chronic response of the liver to natural or manufactured compounds. Up to 20%-40% of DILI patients present with cholestatic and / or mixed hepatocellular / cholestatic patterns of injury. The samples assayed use a cell structure associated with the liver that follows cholestasis into hepatocytes in vitro. The type of cell structure used is, but is not limited to, for example, HepaRG (RTM) cells, which are terminally differentiated hepatocytes derived from a human hepatic progenitor cell line that retains many of the properties of primary human hepatocytes.

[0112] HepaRG cells are immortalized cell lines with 4 main features: 1-Full array of primary human hepatocyte functions, responses and regulatory pathways, including: Phase I and II, and transporter activity consistent with that found within primary human hepatocyte populations; 2-Forming bile canaliculi; 3-Potential to express major properties of stem cells; 4-High plasticity and full trans-differentiation capacity. Cells can be used directly in experimental assay plates after 7 days of maturation. HepaRG cells can form cell spheroids that mimic or simulate one or more cellular processes of the liver, which can be stained and / or fluoresced and imaged for downstream analysis during in vitro microscopy assays.

[0113] Although the toxicity prediction system is described herein with reference to liver cell structures / spheroids (e.g., HepaRG cell lines) and DILI, this is by way of example only and the present invention is not limited thereto, and it will be appreciated by those skilled in the art that the toxicity prediction system (or non-toxic or efficacy prediction system) can be trained and applied to any type of cell structure, any cell line, cell array, and / or cell sample well to which a compound has been added that is associated with any organ and / or any related disease. For example, the toxicity prediction system (or non-toxic / efficacy prediction system) can be applied to cell structure samples that mimic organs, such as, but not limited to, for example, lungs, skin, kidneys, pancreas, liver, cardiac cell structures / hearts, neural cell structures, and / or any other organ of a subject or patient. Such cell structures can be used in samples to which compounds are applied in in vitro microscopy assays, etc., and toxicity is automatically analyzed by a toxicity prediction system.

[0114] FIG2a illustrates an example toxicity prediction pipeline 200 for automatically predicting the toxicity of one or more compounds applied to multiple samples of a cell structure in an in vitro microscopy assay. The toxicity prediction pipeline 200 includes the in vitro HTS microscopy assay system 102 and the quality control imaging analysis system 104 of FIG1a. The downstream analysis system 106 of FIG1a has been modified to form the toxicity prediction system 202. Although the toxicity prediction system 202 is described, this is only an example and the present invention is not limited thereto, and those skilled in the art should understand that other types of prediction systems, such as but not limited to, for example, a non-toxicity prediction system or an efficacy prediction system, can be implemented based on the concepts and DL / ML models described with respect to the toxicity prediction system 202.

[0115] Referring to FIG. 2a, the in vitro HTS assay system 102 is configured to acquire a sample set of cell structures 102a and apply one or more compounds or reagents 102b to the sample set of cell structures 102a to be input into a well sample set in a microscopy assay plate 102c for HTS staining and microscopy assay imaging 102d. In the HTS staining and microscopy assay imaging 102d, the sample wells can be treated / stained with fluorescent reagents / compounds to emphasize the cell structure of each sample. For example, the treated / stained cell structure (such as, for example, spheroid structure, vesicle and / or cell nucleus) can be imaged by a microscopy imaging system. Therefore, the in vitro HTS microscopy assay system 102 is configured to output a well image sample set 103. Each well sample image in the set corresponds to the cell structure in the well sample set measured and each sample to which the compound is applied.

[0116] The mass imaging analysis system 104 can be configured as described with reference to Figures 1a to 1g. In this example, the mass imaging analysis system 104 is configured to receive the well sample image set and perform image preprocessing and / or analysis to identify which samples from the well sample set are feasible for further downstream analysis, such as, for example, toxicity prediction of compounds applied to samples of the well sample set. One or more image processing and / or machine learning algorithms can be applied to identify the feasibility of each well sample image based on any detected image artifacts and / or imaging defects, etc., and / or to enhance or emphasize the cell structure of interest within each feasible well sample image in the well sample set. Therefore, the imaging analysis system 104 can output a feasible well sample image set 105 for further downstream analysis. In essence, the feasible well sample image set 105 can be any suitable image set of the cell structure of the well sample set from the assay, which fully describes the cell structure of the sample for automatic analysis.

[0117] The toxicity prediction system 202 is configured to receive a cell structure image set 105 from a plurality of cell structure samples to which one or more compounds are applied during an in vitro microscopy assay. In this case, the received sample image set 105 may be a feasible well sample image set 105 output from the quality control imaging analysis system 104.

[0118] In the toxicity prediction system 202, the received sample image sets 105 can each be input into a deep learning (DL) toxicity prediction model 202a-202c, which is configured to predict the toxicity of each corresponding compound applied in the assay from the in vitro HTS microscopy assay system 102. The DL toxicity prediction models 202a-202c can be based on any one or more DL modeling techniques / algorithms and / or machine learning (ML) techniques / algorithms that have been used to train DL toxicity models to identify or predict whether each of the received sample images 105 indicates toxicity, even when the compound has not yet been applied to the cell structure of one or more of the well samples. The one or more DL / ML techniques / algorithms can be based on supervised ML, unsupervised ML, and / or semi-supervised ML algorithms, etc. However, for the task of training the DL toxicity prediction models 202a-202c, it has been found that supervised learning is difficult due to the limited number of labeled training data sets associated with cell structures that indicate toxicity or not, depending on whether the compound has been applied. Therefore, the combined supervised / unsupervised DL model training framework may be used to train one or more component models of the DL toxicity prediction models 202a - 202c .

[0119] For example, in the example of FIG. 2a, the DL toxicity prediction model 202a-202c may include a trained machine learning (ML) phenotypic feature extraction (FE) model 202a for extracting phenotypic features of a cell structure from each received sample image of the received sample image set 105. Supervised learning may be used to train the ML phenotypic FE model 202a. The ML phenotypic FE model 202a may be based on, but not limited to, for example, a neural network (NN) classifier that is trained using supervised training of readily available labeled / annotated training data sets to classify images of cells, organoids, spheroids, cell structures, etc. (e.g., classifying images of cell structures to determine whether cells are cancer cells or tumor cells). The NN classifier may be based on any NN structure, such as, but not limited to, for example, a feedforward NN (FNN), a recursive NN (RNN), an artificial NN (ANN), a convolutional NN (CNN), any other type of NN, modifications thereof, or combinations thereof. Before the output classification layer or SoftMax output of the NN classifier, the phenotypic representation of the cell structure may be embedded through the high-dimensional output of one of the hidden layers or complete layers of the NN classifier. The NN classifier is configured to output an embedding of the phenotypic representation of the cell structure from the hidden layer or the complete layer, rather than outputting a classification.

[0120] Typically, the phenotypic representation embedding of the input image samples from the received image sample set 105 is a high-dimensional representation of the phenotypic features (e.g., for a CNN-type NN classifier / model, the dimensionality may be on the order of 2024 or more). The trained NN classifier may be used to output a high-dimensional phenotypic feature representation 204 of the cell structure of each input well sample to which the compound was applied in the sample image set 105. As an example, the neural network classifier may be based on a convolutional neural network (CNN) for classifying images of cell structures (e.g., classifying images of cell structures to determine whether a cell is a cancer cell), wherein, once trained, one of the final complete layers of the CNN may be used as the phenotypic feature representation of each image sample in the received image sample set 105 input thereto.

[0121] The ML phenotype embedding model 202a is coupled to a trained ML low-dimensional (LD) embedding model 202b, which is trained and configured to optimally embed a high-dimensional phenotypic representation 204 of a predicted phenotypic feature of each received sample image of the received image sample set 105 into a lower-dimensional phenotypic embedding 205 in a lower-dimensional space for further analysis, wherein the lower-dimensional phenotypic feature embedding 205 includes those phenotypic features associated with toxicity. Rather than using supervised training on the ML LD embedding model 202b, one or more unsupervised DL / ML techniques / algorithms are used to ensure that the trained ML LED embedding model 202b outputs an LD phenotypic embedding 205 that represents phenotypic features associated with toxicity of the cell structure. The unsupervised DL / ML techniques / algorithms may be based on at least a clustering algorithm or a dimensionality reduction algorithm, such as, but not limited to, for example, a support vector machine (SVM), a uniform manifold approximation and projection (UMAP), or a t-distributed stochastic neighbor embedding (t-SNE) type algorithm, a combination thereof, a modification thereof, and the like.

[0122] For example, the ML LD embedding model 202b can be trained using negative control samples and positive control samples included on the assay plate 102c in an in vitro microscopy assay. That is, the assay plate 102c can include a negative control sample group and a positive control sample group, the negative control sample group can be a first set of wells of the assay plate 102c, wherein the sample is not perturbed or has no compound applied thereto, and the positive control sample group can be a second set of wells of the assay plate 102c, wherein the sample is perturbed by a known compound with known toxicity. The assay plate 102c can also include a third set of wells of the assay plate 102c, the third set of wells including a cell structure sample with a compound applied thereto. Therefore, the negative image group and the control image group associated with the negative control sample and the positive control sample can be used to train the ML LD embedding model 202b in an unsupervised manner.

[0123] For example, a negative image group and a control image group are input to an ML phenotype embedding model 202a, which outputs a corresponding negative high-dimensional phenotype representation and a positive high-dimensional phenotype representation. Then, an iterative optimization of the parameters of a UMAP or t-SNE algorithm can be performed on the high-dimensional negative and positive control phenotype representations output from the ML phenotype embedding model 202a, wherein the difference between the resulting LD negative and positive control phenotype representation embeddings is maximized. The resulting optimized parameters can be used with the UMAP or t-SNE algorithm in the ML LD embedding model 202b for dimensionality reduction of a high-dimensional phenotype representation 205 corresponding to a third set of well samples. The LD phenotype representation 205 can be output for comparison with a negative control (NC) LD phenotype representation using a suitable distance or similarity metric used by the trained ML distance model 202c.

[0124] The ML LD embedding model 202b is coupled to a trained ML distance / prediction model 202c / 202d, which is configured to perform a distance or similarity metric comparison / estimation 206 between the LD phenotype representation of one of the image samples of the well and the negative control LD phenotype representation using a suitable distance or similarity metric associated with the LD space of the LD phenotype representation. The ML distance model 202c can output data representing each distance comparison estimate 206, which is input to the ML prediction unit 202d to predict toxicity. It should be noted that the LD space embedded in the LD phenotype representation output by the ML LD embedding model 202b can still be considered a high-dimensional space that cannot be reliably used by the Euclidean distance metric or similarity metric decomposition. Therefore, a high-dimensional distance metric or similarity metric can be used based on, but not limited to, for example, the Wasserstein distance and / or any other high-dimensional distance metric or similarity metric. The ML distance / prediction model 202c / 202d outputs a toxicity prediction as a probability based on a distance comparison 206 (eg, Wasserstein distance comparison) between each LD phenotype representation of the sample and the negative control LD phenotype representation.

[0125] FIG2b shows an example toxicity prediction process 210 for predicting the toxicity of one or more compounds of multiple samples of cell structures in the in vitro microscopy assay pipeline 200 of FIG2a by the toxicity prediction system 202. The toxicity prediction process 210 includes the following steps: In step 212, a set of images associated with multiple samples is received based on the output of the in vitro microscopy assay. Each sample image includes image data that fully describes the cell structure of the associated sample for automatic processing and analysis. In step 214, each image in the image set is input to a first ML model, which is configured to predict the phenotypic characteristics of the cell structure within the sample associated with each image. In step 216, each of the predicted phenotypic features associated with each sample is input to a second ML model, which is configured to predict the lower dimensional phenotypic feature embedding of each sample. In step 218, the distance between the lower dimensional phenotypic feature embedding of each sample and the lower dimensional phenotypic feature embedding of the sample (e.g., a negative control sample) to which a compound with known toxicity is applied is compared. In step 220, based on the comparison, an indication (eg, probability) of the toxicity of each sample and the compound applied to the sample is output for each sample.

[0126] Although FIG. 2 b shows an example toxicity prediction process 210 for predicting the toxicity of one or more compounds applied to multiple samples of cell structures in the in vitro microscopy assay pipeline 200 of FIG. 2 a, this is for example only and the present invention is not limited thereto, and one skilled in the art will appreciate that the method can be applied to assays that require other downstream analyses 106, which can include, but are not limited to, for example, non-toxicity or efficacy analyses configured to predict the non-toxicity or efficacy of one or more compounds applied to multiple feasible samples of cell structures in an in vitro microscopy assay. For example, process 210 can be modified for use in other downstream prediction systems 106 of FIG. 1 a, which can be configured to predict the non-toxicity, efficacy, and / or other characteristics of one or more compounds applied to multiple samples of cell structures in the in vitro microscopy pipeline 200 of FIG. 2 a. For example, toxicity prediction process 210 can be modified to a non-toxicity or efficacy prediction process based on the following modifications to steps 212-220 of process 210. Step 212 may be further modified or configured to receive an image set associated with a plurality of samples to which one or more compounds associated with non-toxicity and / or efficacy analysis are applied. Step 214 may be further modified to input each image in the image set into a first ML model configured to predict phenotypic features of cell structures within the sample associated with each image. Step 216 may be further modified to input each of the predicted phenotypic features associated with each sample into a second ML model configured to predict a lower dimensional phenotypic feature embedding for each sample, the second ML model having been trained on non-toxicity / efficacy positive / negative controls. Step 218 may be further modified to compare the distance between the lower dimensional phenotypic feature embedding of each sample and the lower dimensional phenotypic feature embedding of a sample to which a compound with known non-toxicity or efficacy is applied. Step 222 may be further modified to output, for each sample, an indication of the non-toxicity or efficacy of the compound applied to the sample based on the comparison.

[0127] Figure 3An example neural network classifier 300 in the ML phenotypic FE model 202a for the toxicity prediction system 202 of FIG. 2a is shown. The NN classifier 300 is configured to predict or output a phenotypic representation / embedding from an image of a cell structure in a sample determined by microscopy of FIG. 2a. The NN classifier 300 is based on a CNN-type architecture that includes a first portion of a CNN-type network 302, followed by one or more fully connected layers 304a-304n and a classifier output layer 306. One or more of the outputs 305a-305n of the fully connected layers 304a-304n may be tapped and / or selected 308 for output as a phenotypic representation embedding 310 associated with an input sample image representing a cell structure. As an example, the NN classifier 300 may be a 50-layer RESNET CNN architecture (e.g., RESNET50 (RTM) or VGG (RTM) etc.) with multiple fully connected layers.

[0128] The NN classifier 300 is trained on image data containing cell structures relevant to the classification task, wherein relevant phenotypic information of the cell structures can be extracted from the output of one of the hidden layers or fully connected layers of the trained NN classifier 300. The NN classifier 300 can be pre-trained for the classification task using labeled image data associated with cell types / structures. For example, the NN classifier 300 can be trained using a labeled training image dataset to classify images as, for example, cancerous or non-cancerous, identification of cells / cell structures, and / or other diseases that affect cell function / structure.

[0129] Essentially, when trained for a classification task, the layers within the CNN architecture of the NN classifier 300 begin to "recognize" cellular primitives, cell agglomerations, how cells form tissue regions, etc., to identify larger cellular macrostructures, which can be used to form a high-dimensional phenotypic representation of the cellular structure present in the input image embedded in the cellular structure sample in the complete layers 304a-304n. The classification task or output 306 is only used to train the NN classifier 300. As another example, the NN classifier 300 can be trained on labeled cell image data from ImageNet and then fine-tuned using a limited number of training image data items associated with cell data from microscopy assays, etc. and / or stained / fluorescent cell images.

[0130] The NN classifier 300 architecture can be configured to have multiple fully connected layers 304a-304n before outputting the classification 306. One of the fully connected layers 304a-304n is selected from all of the fully connected layers 304a-304n based on the layer determined to focus on the cellular structure in terms of deep representation. As an option, the toxicity prediction system 202 can perform automatic optimization on each of the ML models 202a-202c to determine which of the fully connected layers 304a-304n of the trained NN classifier 300 provides the best high-dimensional phenotypic representation output that highlights the toxic effect or toxicity in the cellular structure.

[0131] 4a is an example assay plate 400 having a negative control sample well group 402 and a positive control sample well group 404, wherein the samples therein are imaged during an in vitro microscopy assay for use in training a deep learning (DL) toxicity model 202. The assay plate 400 also includes a test sample well group 406, wherein the samples therein are imaged during an in vitro microscopy assay for input into a trained DL toxicity model for predicting the toxicity of any compound applied to the test sample well group 406. The negative control sample well group 402 includes samples that have not been perturbed by the test compound or samples that have been perturbed by a non-toxic compound. For example, in an in vitro microscopy assay using hepatic spheroids, dimethyl sulfoxide (DMSO) can be applied to NC samples (e.g., 0.6% DMSO). The positive control sample well group 404 includes samples that have been perturbed by a compound with known toxicity associated with the cell structure used in the sample. For example, in an in vitro microscopy assay using hepatic spheroids, PC samples can be perturbed with chlorpromazine (CPZ) at concentrations that produce a toxic effect (eg, 400 μM).

[0132] The DL toxicity model 202 is trained and configured to first extract phenotypic features of the cell structure of each sample in the sample image, and then estimate the phenotypic distance between the extracted phenotypic features and the phenotypic features associated with the samples in the negative control sample well group 402. Microscopy images of negative control samples and positive control samples in the negative control sample well group 402 and the positive control sample well group 404, respectively, are used together with unsupervised DL / ML techniques to train the DL toxicity model 202 to estimate the toxicity of the compound of each sample using the phenotypic distance between the extracted phenotypic features and the phenotypic features associated with the samples in the negative control sample well group 402. For example, the negative control samples are used to establish the average and deep phenotypic representation embedding of the negative control samples (e.g., the first 20 wells). Once the average and deep phenotypic representation embedding of the negative control samples have been established, this can be used as a reference for determining the distance of the low-dimensional embedding of subsequent test samples.

[0133] In this example, the assay plate 400 has an array of wells consisting of rows and columns of wells (e.g., columns AP and rows 01-24) that are mapped to negative control sample well groups 402, positive control sample well groups 404, and test sample well groups 406. In this example, the assay plate 400 has an array consisting of 16 columns and 24 rows of wells, or a total of 484 wells. Although the assay plate 400 shows the negative control well group 402 as being mapped to columns AP and rows 1-8, the positive control well group 404 as being mapped to columns AP and rows 20-24, and the test sample well group 406 as being mapped to columns AP and rows 9-19, this is by way of example only and the present invention is not limited thereto, and those skilled in the art will appreciate that any mapping may be defined to specify the mapping between the sample wells of the assay plate 400 and the negative control sample wells 402, the positive control sample wells 404, and the test sample wells 406 groups. The mapping of samples to wells in the assay plate 400 can be automatically defined by the software management system of the HTS in vitro microscopy assay system 102 based on the number of samples required in each group of negative control sample wells 402, positive control sample wells 404, and test sample wells 406. For example, a sample in a test sample well group can have one or more compounds at different concentrations applied to the sample. The HTS in vitro microscopy assay system 102 can be programmed to map the type and concentration of the compound onto the assay plate 400, depending on the number of compounds to be tested, the number of different concentrations, the number of replicates, etc.

[0134] FIG4b shows an example unsupervised training process 410 for training a deep learning toxicity model, which includes the ML low-dimensional (LD) embedding model 202b and the ML distance model 202c of the toxicity prediction system 202 of FIG2a. The training DL toxicity model uses the negative control sample well 402 and the positive control sample well 404 of the sample of FIG4a. The unsupervised training process 410 includes the following steps:

[0135] In step 412, the training process 410 receives negative control (NC) and positive control (PC) phenotypic representation embeddings associated with the NC and PC samples corresponding to the NC sample well group 402 and the PC sample well group 406 of the plate 400. Images of the NC and PC samples corresponding to the NC sample well group 402 and the PC sample well group 406 of the plate 400 (e.g., feasible images from the quality control image analysis system 104) can be input to the ML phenotypic feature extraction model 202a of FIG. 2a, which can be based on Figure 3 The NN classifier 300 is pre-trained as described with reference to Figures 2a-3. In view of this, the ML phenotypic feature extraction model 202a outputs the corresponding NC and PC phenotypic representation embeddings 204. The NC and PC phenotypic representation embeddings are high-dimensional embeddings (e.g., 2024 elements per image).

[0136] In step 413, joint unsupervised training is performed on the ML LD embedding model 202b and the ML distance / prediction model 202c / 202d based on the following steps: In step 414, unsupervised training is performed on the ML LD embedding model 202b to estimate the LD phenotype embedding 205 representing the toxicity signature associated with the NC and PC phenotype representation embeddings. The unsupervised training can be performed by iteratively optimizing the ML LD embedding model 202b within a parameter set range associated with the ML technique (e.g., UMAP or t-SNE) used to configure the ML LD embedding model 202b. The criterion for the ML technique used to adjust or generate the ML LD embedding model 202b is to find a parameter subset that maximizes the difference between the NC embedding and the PC embedding within the parameter set range using the input high-dimensional NC and PC phenotype representation embeddings. In each iteration, the parameters are subsetted from the parameter set range and applied to the ML technique (e.g., UMAP or t-SNE), which adjusts or generates the ML LD embedding model 202b to output LD NC and PC phenotype embeddings.

[0137] As an example, the input high-dimensional NC and PC phenotype representation embeddings 204 output from the ML phenotype FE model 202a are high-dimensional vectors for each image (e.g., each embedding can be 2024 elements per image). The ML phenotype FE model 202a has not been trained specifically for toxicity prediction, but rather for identifying biological structures or phenotypic representations of cellular structures within each image of the sample. Therefore, the resulting high-dimensional NC and PC phenotype representation embeddings describe the phenotypic representation of each corresponding cellular structure as much as possible within the dimension (e.g., 2024). The UMAP algorithm can be used to generate the ML LD embedding model 202b to not only reduce the dimensionality of the high-dimensional NC and PC phenotype representations, but also ensure that the phenotypic representations associated with toxicity are retained or concentrated within the resulting LD NC and PC embedding vectors 205 (e.g., each embedding can be 64 elements per image) output from the MLLD embedding model 202b. This is performed by the UMAP algorithm finding the optimal parameter set that maximizes the difference between the high-dimensional NC phenotype representation embedding and the PC phenotype representation embedding within the range of the parameter set. The ML LD embedding model 202b is trained to find relevant biological structures associated with toxicity within each image in an unbiased manner without knowing the cell type or its application. The ML LE embedding model 202b can use the LD embedding (e.g., 64 dimensions in the embedding) to represent the entire negative control group by taking a simple average profile of all negative controls. This can be used by the ML distance / prediction model 202c / 202d to determine a toxicity prediction based on estimating the distance or similarity between the LD embedding of the test sample from the test sample well group 406 and the average LD NC embedding of the NC samples from the NC sample well group 402 using a distance or similarity metric.

[0138] In step 415, the ML LD embedding model 202b outputs NC and PC LD phenotypic embeddings 205, wherein each embedding includes phenotypic information of the corresponding cellular structure associated with toxicity. The NC and PC LD phenotypic embeddings are used as input to step 416 for training the ML distance model 202c.

[0139] In step 416, unsupervised training is performed on the ML distance model 202c to estimate a suitable distance or similarity based on a high-dimensional vector metric (e.g., Wasserstein distance) for maximizing the distance between the NC LD phenotype embedding and the PC LD phenotype embedding. This means that the trained ML distance model 202c can determine the distance 206 between the NC LD phenotype embedding and the LD phenotype embedding of the test sample for determining a probability or indication of whether the compound associated with each test sample is toxic. The unsupervised training can be performed by iteratively optimizing the ML distance model 202c within a parameter set associated with the ML distance algorithm / technique used to configure the ML distance model 202c (e.g., Wasserstein distance algorithm, such as Earth Mover's distance and / or Sinkhorn distance algorithm). The criterion for the ML distance algorithm / technique used to adjust or generate the ML distance model 202c is to find a parameter subset of the ML distance algorithm that maximizes the distance between the NC LD embedding set and the PC embedding set, but minimizes the distance between embeddings within the NC LD embedding set and minimizes the distance between embeddings within the PC LD embedding set. In each iteration, a parameter subset is selected from the range of parameter sets for the ML distance algorithm and applied to the ML distance algorithm (e.g., the bulldozer moving distance algorithm or the Sinkhorn algorithm), which adjusts or generates the ML distance model 202c / 202d for outputting a toxicity indication between the LD phenotypic embedding and the NC LD embedding set.

[0140] In step 417, it is determined whether an optimal distance between the NC LD embedding set and the PC LD embedding set has been obtained for the range of parameter sets for the ML distance model 202c. This may also include determining whether the difference between the NCLD embedding and the PC LD embedding has been maximized based on the parameter sets tested to date. If this is the case or if there are no more parameter combinations in the parameter set of the ML LD embedding model 202b or the ML distance model 202c that can be used, the process 410 proceeds to step 419. Otherwise, the process proceeds to step 418 for further parameter selection and adjustment / training of the ML LD embedding model 202b and / or the ML distance model 202c.

[0141] In step 418, additional parameters are selected from the parameter set associated with the ML LD embedding model 202b for further training thereof, wherein the process 410 proceeds to step 414. Similarly, additional parameters may be selected from the parameter set associated with the ML distance model 202c for further training thereof, wherein the process 410 proceeds to step 416.

[0142] In step 419, the ML LD embedding model 206b selects a subset of parameters within the parameter set associated with the ML LD embedding model 202b that maximizes the difference between the NC high-dimensional phenotypic embedding and the PC high-dimensional phenotypic embedding for outputting an LD phenotypic embedding associated with the high-dimensional phenotypic embedding of the test samples corresponding to the test sample well group 406. Similarly, the ML distance / prediction model 202c / 202d selects a subset of parameters within the parameter set associated with the ML distance model 202c that maximizes the distance (e.g., Wasserstein distance) between the NC LD phenotypic embedding set and the PC LD phenotypic embedding set, but also minimizes the distance within each of the NC and PC LD phenotypic embedding sets, for outputting a toxicity estimate based on the distances 206 between the LD phenotypic embeddings of the test samples corresponding to the test sample well group 406 and the NC phenotypic embedding set or the average NC phenotypic embedding.

[0143] FIG4c shows a toxicity prediction process 420 for predicting the toxicity of a compound using the trained deep learning toxicity model of FIG4b. From step 419 of the toxicity training process 410 of FIG4b, the ML LD embedding model 202b and the ML distance / prediction model 202c / 202d are configured based on the output selected parameter set 410. The DL toxicity model of FIG4b includes the ML LD embedding model 202b trained as described in reference to FIG4b and the ML distance model 202c trained as described in reference to FIG4b. The toxicity prediction system 202 includes the ML phenotype model 202a as described in reference to FIG2a to 3 and the DL toxicity model of FIG4b, which is used to predict the toxicity of one or more compounds of multiple test samples applied to a cell structure in an in vitro microscopy assay. The multiple test samples in the sample test well group 406 of FIG4a can be captured by a microscopy imager during an in vitro microscopy assay. During an in vitro microscopy assay, a collection of one or more compounds can be applied to the test sample, i.e., one compound per test sample. These can be processed and / or enhanced using the quality control imaging system 104 of Figures 1a-1g. The resulting images of the test sample can then be input into the toxicity prediction system 202 for predicting the toxicity of the compound applied to the test sample. The toxicity prediction process 420 includes the following steps:

[0144] In step 421, a set of images associated with one or more test samples to which a compound is applied is received from a sample test well group 406 of an in vitro microscopy assay. The compound may have known or unknown toxicity to the cell structure within the corresponding test sample. Each image of the test sample includes image data that fully describes the cell structure of the relevant test sample to which the compound is applied for automatic processing and analysis. In step 422, each image in the image set is input to the trained ML phenotypic feature extraction model 202a or 300 of Figure 2a or 3. For each input image of the test sample, the trained ML phenotypic feature extraction model 202a or 300 outputs a high-dimensional phenotypic representation 204 of the cell structure of the test sample located within the input image of the test sample. The output high-dimensional phenotypic representation 204 of each test sample is then applied to the trained ML LD embedding model 202b. In step 433, each of the high-dimensional phenotypic representations 204 of each test sample is input to the trained ML LD embedding model 202b for outputting a low-dimensional phenotypic embedding 205 of each test sample. In step 424, the LD phenotype embedding 205 of each test sample is passed to the trained ML distance model 202c.

[0145] In step 425, the ML distance model 202c receives each LD phenotype embedding 205 for each test sample and outputs a distance estimate 206 between the LD phenotype embedding of each test sample and the set of NC LD phenotype embeddings. In step 425a, the ML distance model 202c can be configured to output a distance 206 or similarity estimate 206 between the LD phenotype embedding of each test sample and the average NCLD embedding of the set of NC LD phenotype embeddings. This can include, in step 425b, comparing the distance between the LD phenotype embedding of the test sample and the average NC LD embedding of the set of NC LD embeddings. From step 425, the ML distance / prediction model 202c / 202d can output an indication or probability associated with the distance of each LD phenotype embedding in the LD phenotype embedding to the NC LD embedding (i.e., determining how far the LD phenotype embedding is from the non-toxic NC LD phenotype embedding). In step 426, based on the comparison of the ML distance model 202c, an indication (e.g., probability) of the toxicity of each test sample and the compound applied to the test sample is output for each test sample.

[0146] 5 a is a schematic diagram illustrating another example assay plate 500 having a negative control sample group 502 and a positive control sample group 504 for training a deep learning (DL) toxicity model of the toxicity prediction system 202 and the model as described with reference to FIGS. 2 a , 3 , and 4 a - 4 b , and a test sample group 506 for input to the trained DL toxicity model of the toxicity prediction system 202 .

[0147] In this example, the HepaRG liver cell line is used to determine the toxicity of the compound in the liver. Each of the sample wells of the sample plate 500 is filled with a HepaRG cell structure, wherein the in vitro microassay system 102 of FIG. 2a is configured to use a specific fluorescent substrate (e.g., carboxyl-DCFDA (5-(and-6)-carboxyl-2',7'-dichlorofluorescein diacetate) or CDFDA) to evaluate the cellular cholestatic effect of the compound. CDFDA is a reagent that passively diffuses into the cell. It is a fluorescent substrate for the imager to capture the accumulation of CDF in the bile ductules in the image of each sample. This enables the assessment of whether the compound causes cholestasis, which has been observed to occur when the bile ductules of the cell structure disappear in the image. However, the toxicity prediction system 202 performs further unbiased processing to take into account other unobserved changes in the cell structure when determining the toxicity of the compound.

[0148] The samples with HepaRG cell structures in the negative control sample set 502 only have a buffer (e.g., DMSO) applied to them that has no toxic effects. The samples with HepaRG cell structures in the positive control sample set 504 have a reference compound with a known toxic effect (e.g., CPZ at 60 micromolar), which is known to be toxic to hepatocytes and trigger cholestasis. The samples with HepaRG cell structures in the test sample set 506 have a variety of different compounds applied at different concentrations and replicates. In this example, each compound can be represented by 8 doses in the assay plate 500, where each dose is represented by 3 replicates in the assay plate 500. This is to ensure that there are enough replicates so that after quality control is applied to the images of the test samples, there will be at least one replicate for each compound and each dose with a viable image of the test sample for further downstream analysis.

[0149] Capture the image of each sample well in the sample well of the assay plate 500, and evaluate the feasibility associated with further downstream analysis by the toxicity prediction system 202. This can be performed by the quality control image system 104 of Figures 1a to 1g or 2a. For example, the quality control model can be trained to automatically evaluate the feasibility of each well sample by classifying the relevant images of the test sample in the well. Each well sample is classified with the probability of indicating whether the well sample is of good quality or poor quality. The higher the probability value, the better the quality, and the lower the probability value, the worse the quality. A threshold probability can be used to determine a feasible high-quality sample well. In this example, a probability value of about 0.1 is determined to generate a feasible sample image that can be passed to the toxicity prediction system 202 for toxicity training and / or toxicity prediction. The light gray shaded area (e.g., the image of the sample in the well 502a-502c, 504a, 506a-506c) represents the sample in the NC, PC and test sample well indicating the feasible sample for further downstream analysis. The darkest gray shaded areas (e.g., images of samples in wells 502d, 506b, and 504d) represent samples in NC, PC, and test sample wells that indicate infeasible samples that have too many artifacts for analysis and can be discarded. Therefore, a feasible sample image set from each of the NC group 502, PC group 504, and test sample group 506 can be used for further downstream analysis. The feasible image sets from the NC sample group 502 and the PC sample group 504 are used to train a DL toxicity model in an unsupervised manner, the DL toxicity model including training the ML LD embedding model 202b and ML distance / prediction model 202c / 202d of FIG. 2a, as described with reference to FIGS. 2a to 4b. The toxicity prediction system includes an ML phenotypic feature extraction model 202a and a DL toxicity model, the DL toxicity model including the ML LD embedding model 202b and ML distance / prediction model 202c / 202d of FIG. 2a.

[0150] Figure 5b shows an example distance matrix of negative and positive control samples of a trained DL toxicity model trained based on the unsupervised training process of Figures 2a and 4a using feasible NC and PC samples from the NC sample set 502 and PC sample set 506 on the assay plate 500 of Figure 5a. Once the feasible images of NC and PC samples have been passed through the ML phenotypic feature extraction model 202a or 300 of Figures 2a-3, the output set of NC and PC high-dimensional phenotypic embeddings is used as input to the UMAP algorithm to train and generate the ML LD embedding model 202b, as described with reference to Figure 4b.

[0151] For example, a grid search is used to iteratively optimize a parameter set defining hyperparameters of the UMAP algorithm, wherein different hyperparameter values ​​from the parameter set are iteratively applied to determine an optimal combination of UMAP parameters that maximizes the difference between the NC low-dimensional phenotype embedding set and the PC low-dimensional embedding set as much as possible. Assuming that the LD phenotype embedding is still high-dimensional (e.g., 64 elements), Euclidean, Manhattan, and other standard distance metrics cannot be applied, so the Wasserstein distance is used instead.

[0152] As described with reference to Figures 2a and 4b, the output set of NC and PC LD phenotype embeddings from the ML LD embedding model 202b based on the UMAP algorithm is applied to the Sinkhorn algorithm for training and generating the ML distance model 202c, as described with reference to Figure 4b. The Wasserstein distance is optimized by iteratively performing a grid search within the parameter set to find hyperparameters for the Sinkhorn algorithm that maximize the estimated Wasserstein distance between the set of NC LD phenotype embeddings and the set of PC LD phenotype embeddings. The optimized hyperparameters output by the Sinkhorn algorithm can then be used by the ML distance / prediction model 202c / 202d for other test sample LD embeddings to enable comparison of distances with negative control LD embeddings.

[0153] The matrix distance graph in FIG5b is a mapping of distances when the toxicity prediction model 202 has been calibrated to use feasible input sample images from the assay plate 500 of FIG5a. Columns 0 to 9 and rows 0 to 9 of the matrix graph represent the distances between the NC LD embeddings associated with the feasible image samples of the NC sample well 502 of FIG5a. Columns 10-19 and rows 10-19 of the matrix graph represent the distances between the PC LD embeddings associated with the feasible image samples of the PC sample well 504 of FIG5a. Clearly, the DL toxicity model (i.e., the ML LD embedding model 202b and the ML distance / prediction model 202c / d) has been trained such that the distances between the NC LD embeddings have a dark gray shaded area 512 indicating the minimum distance therebetween. Similarly, for the PC LD embeddings, it also has a dark gray shaded area 514 indicating the minimum distance therebetween. Likewise, for this example, the light gray shaded area 516 indicates that the distances between the NC LD embedding set and the PC LD embedding set have been maximized (as much as possible). Assume that the phenotypic distance in region 516 is maximized, indicating that the corresponding PC sample has a toxic effect on samples of the HepaRG cell line. This clearly indicates that the DL toxicity model of the toxicity prediction system 202 has been calibrated and can be used to test the toxicity of a series of compounds associated with the HepaRG cell line.

[0154] FIG5c shows another example distance matrix 520 of negative and positive control samples and test samples used to predict the toxicity of compounds of test samples using the trained DL toxicity model 202b-202d of FIG5b of the toxicity prediction system 202. In addition to the feasible NC and PC sample images output from the in vitro microassay represented by the assay plate 500 of FIG5a, multiple feasible test sample images are processed to predict the toxicity of compounds applied to the test samples. In the matrix 520, columns 0 to 95 and 288 to 311 and rows 0 to 95 and 288-311 represent feasible images of NC samples used from the plate 500, columns 312 to 335 and rows 312-335 represent feasible images of PC samples used from the plate 500 to train the DL toxicity model 202b-202d of FIG5b, and columns 96 to 287 represent feasible images of test samples used from the plate 500 for testing. As can be seen, the dark grey area 522 in columns 0 to 95 and rows 0 to 95 of the distance matrix represents the distance between the NC LD embeddings, with the minimum distance between them. There are some false positives given light grey shading, which does not affect the performance of toxicity prediction. Similarly, the dark grey area 524 in columns 312 to 335 and rows 312 to 335 of the distance matrix 520 represents the distance between the PC LD embeddings, which also has the minimum distance between them. Similarly, the light grey shaded area 526 of the distance matrix represents the distance between the NC LD embedding set and the test sample LD embedding set, which does not represent the minimum distance from the NC LD embedding, but rather indicates a larger distance that has a higher toxic effect on the HepaRG cell line sample used in most (if not all) test samples.

[0155] Figure 5d shows an example of a conventional toxicity prediction method used for a set of 14 compounds (e.g., compounds A, B, C, D, E, F, G, H, I, J, K, L, M and N) applied to a sample of a HepaRG cell structure. In this example, the assay plate has 14 compounds, 8 doses, and each dose has 3 repetitions. It is known that these 14 compounds have toxic effects on the HepaRG cell structure and therefore on the liver. These compounds are tested using commercial software accompanying in vitro microscopy hardware and conventional toxicity analysis. The conventional toxicity analysis is based on the use of standard image analysis tools to analyze data (e.g., Phase 1) and standard toxicity characterization based on vesicle counting (e.g., Phase 2) and researchers' analysis of dose-response (EC50) charts (e.g., Phase 3). In Phase 4, it was found that only 8 compounds (e.g., compounds A, E, F, I, J, K, M and N) had toxic effects, but the conventional toxicity analysis method determined that 6 toxic compounds (e.g., compounds B, C, D, G, H and L) had no toxic effects. However, it is clear that the conventional toxicity analysis workflow does miss compounds that have toxic effects because 6 compounds cannot be detected. When these compounds are applied to the toxicity prediction system 202 and the trained DL toxicity model 202b / 202c / 202d, as described with reference to Figures 5a-5c, it is found that all compounds have toxic effects. Figure 5e is a schematic diagram showing an example of conventional toxicity prediction results of compounds A, B, C, D, E, F, G and I obtained using a conventional method for predicting toxic compounds. The bile vesicle count and cell count of the positive control compound (i.e., chlorpromazine) are indicated in the dotted box, and the bile vesicle count and cell count of the negative control compound (i.e., DMSO) are indicated by the solid box. As can be seen, when only the bile vesicle count and cell count are considered, it is difficult to predict whether compounds B, C, D and G have toxic effects. Figure 5f is a schematic diagram showing an example of the toxicity prediction results of the trained DL toxicity model of compounds A, B, C, D, E, F, G and I obtained using the trained DL toxicity model 202b / 202c / 202d. The schematic diagram of Figure 5f is a graphical representation of the novel phenotypic distance metric (y-axis) of the deep learning-based phenotype extracted using a cell-based imaging classification framework (e.g., Resnet50). The phenotypic distance metric of the positive control compound (i.e., chlorpromazine) (dashed box) and the negative control compound (i.e., DMSO (in the solid box)) is shown. The phenotypic distance metric compares each compound with the negative control compound (average phenotypic representation), so the distance (y-axis) between the tested compound (e.g., compound A, B, C, D, E, F, G and I) and the average negative DMSO phenotypic profile is greater.As can be seen, for each tested compound, the distance between the compound phenotype and the reference negative profile (average DMSO) is above the 75th percentile of the DMSO distribution. In addition to the similarity to the positive control compound (chlorpromazine), this difference can also characterize the phenotypic effect as toxicity (the difference from DMSO alone cannot be characterized as toxicity alone). As can be seen, using the trained DL toxicity model 202b / 202c, compounds A, B, C, D, E, F, G and I are more easily predicted to be toxic, and the results can be used to more effectively predict the toxicity of compounds compared to conventional toxicity prediction methods.

[0156] Although the toxicity prediction system 202 has been described with reference to Figures 2a to 5e, this is only an example and the present invention is not limited thereto. Those skilled in the art should understand that, according to application requirements, the methods and processes for training the models 202a-202d of the toxicity prediction system 202 can be modified and / or applied to alternatively predict the non-toxicity of a compound and / or predict the efficacy of a compound and / or predict any other properties of a compound.

[0157] Figure 6a is a schematic diagram of a system / apparatus for performing the methods described herein. The system / apparatus shown is an example of a computing device. Those skilled in the art will appreciate that other types of computing devices / systems may alternatively be used to implement the methods described herein, such as a distributed computing system.

[0158] The device (or system) 600 includes one or more processors 602. The one or more processors control the operation of other components of the system / device 600. For example, the one or more processors 602 may include a general-purpose processor. The one or more processors 602 may be a single-core device or a multi-core device. The one or more processors 602 may include a central processing unit (CPU) or a graphics processing unit (GPU). Alternatively, the one or more processors 602 may include dedicated processing hardware, such as a RISC processor or programmable hardware with embedded firmware. Multiple processors may be included.

[0159] The system / apparatus includes a working or volatile memory 604. One or more processors can access the volatile memory 604 to process data and can control the storage of data in the memory. The volatile memory 604 can include any type of RAM, such as static RAM (SRAM), dynamic RAM (DRAM), or it can include flash memory, such as an SD card.

[0160] The system / apparatus includes non-volatile memory 606. The non-volatile memory 606 stores a set of operating instructions 608 in the form of computer-readable instructions for controlling the operation of the processor 602. The non-volatile memory 606 may be any type of memory, such as read-only memory (ROM), flash memory, or magnetic drive memory.

[0161] The one or more processors 602 are configured to execute operating instructions 608 to cause the system / device to perform any of the methods described herein. The operating instructions 608 may include code related to the hardware components of the system / device 600 (i.e., drivers), as well as code related to the basic operation of the system / device 600. In general, the one or more processors 602 execute one or more instructions in the operating instructions 608 permanently or semi-permanently stored in the non-volatile memory 606, and use the volatile memory 604 to temporarily store data generated during the execution of the operating instructions 608.

[0162] Figure 6b 6 is a schematic diagram of a system 610 for performing the methods described herein. The system 610 shown is an example of a computing device, system, and / or cloud computing system, etc. Those skilled in the art will appreciate that other types of computing devices / systems may alternatively be used to implement the methods described herein, such as distributed computing systems. The system 610 includes a sampling module / unit 612, a first imager module / unit 614, a sample feasibility module / unit 616, and an output module / unit 620, which may be connected together or communicate with each other as needed to implement the methods and / or apparatus / systems as described herein.

[0163] For example, the sampling module 612 can be configured to identify a first sample set that can be used for analysis from a plurality of samples of an assay plate. The imager module 614 can be configured to generate a 2-dimensional (2D) image set for each sample in the first sample set, the 2D image set of each sample comprising a plurality of 2D image slices taken along the z-axis of each sample. The sample feasibility module 216 can be configured to identify a feasible sample set from the 2D image slice set. The output module 218 can be configured to output data representing the feasible sample set for analysis. The data representing the feasible sample set for analysis can be input to, but not limited to, another system 620, such as the toxicity prediction system 202 or configured to predict the toxicity / non-toxicity and / or other characteristics of one or more samples, etc.

[0164] Figure 6c6 is a schematic diagram of another system 620 for performing the methods described herein. The system 620 shown is an example of a computing device, system and / or cloud computing system, etc. Those skilled in the art will appreciate that other types of computing devices / systems may alternatively be used to implement the methods described herein, such as distributed computing systems. The system 620 includes a receiver module / unit 622, a first ML model module / unit 624, a second ML model module / unit 626, a distance comparison module / unit 628, and an output module / unit 630, which may be connected together or communicate with each other as needed to implement the methods and / or apparatus / systems described herein.

[0165] For example, the receiver module 622 may be configured to receive an image set associated with a plurality of samples. The image set may include, but is not limited to, for example, data representing a feasible sample set for analysis output from the system 610. The first ML model module 624 may be configured to input each image in the image set into the first ML model 202a, which is configured to predict the phenotypic features 204 of the cell structure within the sample associated with each image. The second ML model module 626 may be configured to input each of the predicted phenotypic features 204 associated with each sample into the second ML model 202b, which is configured to predict the lower dimensional phenotypic feature embedding 205 of each sample. The distance comparison module 628 may be configured to compare the distance between the lower dimensional phenotypic feature embedding 205 of each sample and the lower dimensional phenotypic feature embedding of the sample with the compound of known toxicity, which may output data representing the comparison 206. The output module 630 may be configured to output, for each sample, an indication of the toxicity of the compound applied to the sample and the compound based on the comparison 206.

[0166] Embodiments of the methods described herein may be implemented in digital electronic circuit systems, integrated circuit systems, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These may include computer program products (such as software stored on, for example, a disk, an optical disk, a memory, a programmable logic device), which include computer readable instructions that, when executed by a computer, for example, Figure 6a , 6b As described in and / or 6c, a computer is enabled to perform one or more methods described herein. Any system feature as described herein may also be provided as a method feature, and vice versa. As used herein, means plus function features may alternatively be expressed with their corresponding structures. Specifically, method aspects may be applied to system aspects, and vice versa.

[0167] In addition, any, some and / or all features in one aspect may be applied to any, some and / or all features in any other aspect in any appropriate combination. It should also be understood that the specific combination of various features described and defined in any aspect of the present invention can be implemented and / or provided and / or used independently.

[0168] While several embodiments have been shown and described, it will be appreciated by those skilled in the art that changes may be made in these embodiments without departing from the principles of the present disclosure, the scope of which is defined in the claims.

Claims

1. A computer-implemented method for identifying viable samples of cellular structures for analysis in an in vitro microscopy assay, the method comprising: include: automatically identifying (111) a first set of samples available for analysis from a plurality of samples on an assay plate; generating (112) a 2-dimensional 2D image set for each sample in the first sample set, the 2D image set for each sample comprising a plurality of 2D image slices captured along a z-axis of each sample; identifying (113) a feasible sample set from the set of 2D image slices; as well as The output (114) represents data of the feasible sample set for analysis as the image set.

2. The computer-implemented method of claim 1, in, Automatically identifying the first set of samples further comprises, for each sample in the plurality of samples: preprocessing (116) the image of each sample; Inputting (117) the pre-processed sample image to a first machine learning (ML) model (124), the first ML model being configured to identify a region of interest including a cell structure of the input sample image; inputting (118) the identified region of interest of the sample image to a second ML model (130) configured to classify whether the sample is analyzable; and The output (119) includes the first sample set of data representing those samples classified as analyzable.

3. The computer-implemented method of claim 2, in, The first ML model (124) is a convolutional neural network (CNN) or other neural network trained to identify regions of interest including cellular structures, and The second ML model (130) is a one-class SVM configured to classify whether the region of interest is analyzable.

4. The computer-implemented method of claim 3, in, Training and configuring the CNN based on a labeled training dataset, wherein the labeled training dataset includes a plurality of images, each of the images being annotated with a label, the label including data indicating whether a cell region of interest exists and / or a location of the region of interest within the image; and The one-of-a-kind SVM configured to classify whether the region of interest is analyzable is trained and configured.

5. The computer-implemented method of any one of claims 1 to 4, in, Identifying the feasible sample set from the set of 2D image slices further comprises, for each sample: identifying (142) a foreground, a background, and a plurality of uncertain feature regions of the cell structure in each of the 2D image slices, wherein the plurality of uncertain feature regions include a plurality of uncertain foreground features and a plurality of uncertain background features; iteratively combining (144) the foreground, background, and the plurality of uncertain feature regions of the 2D image slices to generate a single 2D image of the cell structure; and selecting (146) the sample for the feasible sample set based on the quality of the single 2D image; and Data representing images of feasible samples associated with the set of feasible samples is output (148).

6. The computer-implemented method of claim 5, in, Outputting data representing an image of the feasible sample set further comprises outputting data representing one or more of the group consisting of: A set of 2D images generated for each feasible sample in the feasible sample set; A preprocessed image of each feasible sample in the feasible sample set; A single 2D image generated for each feasible sample, each single 2D image generated based on iteratively combining the 2D image slices of the feasible sample according to the identified foreground, background and uncertainty regions of the 2D image slices of the feasible sample; as well as Any other images captured or processed in relation to the viable sample.

7. A computer-implemented method as claimed in any preceding claim, in, The cell structure includes one or more from the group consisting of: Cell spheroid structure; Vesicles; Organoids; and Any other suitable cell structure.

8. A computer-implemented method as claimed in any preceding claim, in, The in vitro microscopy assay is a high throughput screening in vitro microscopy assay; and The plate includes a plurality of wells, wherein each well contains a sample of the cell structure.

9. The computer-implemented method of any preceding claim, inputting (105) data representing each of the viable samples into a third ML model (202) configured to perform downstream assay analysis on the viable samples to predict an assay analysis result for each of the viable samples, in, The first subset of the samples includes negative controls, the second subset of the samples includes positive controls, and the third subset of the samples includes samples to be analyzed, wherein the third ML model (202) is trained based on the negative controls / the positive controls.

10. The computer-implemented method of claim 9 or 9, wherein the assay analysis comprises at least one item from the group consisting of: Toxicity analysis; Non-toxic analysis; Efficacy analysis; and Any other analysis.

11. The computer-implemented method of claim 10, in, The assay analysis comprises a toxicity assay configured to predict toxicity of one or more compounds applied to a plurality of viable samples of a cell structure in the in vitro microscopy assay, the method comprising: receiving (212) a set of images associated with the plurality of samples; Inputting (214) each image in the image set to a first ML model (202a), the first ML model being configured to predict a phenotypic characteristic (204) of a cellular structure within a sample associated with said each image; inputting (216) each of the predicted phenotypic features associated with each sample into a second ML model (202b) configured to predict a lower dimensional phenotypic feature embedding (205) for each sample; comparing (218) the distance between the lower dimensional phenotypic feature embedding (205) of each sample and the lower dimensional phenotypic feature embedding of the sample to which the compound with known toxicity was applied; and Based on the comparison (206), an indication of the toxicity of each sample and the compound applied to the sample is output (220) for each sample.

12. The computer-implemented method of claim 10, in, The assay analysis comprises a non-toxicity assay or an efficacy assay configured to predict the non-toxicity or efficacy of one or more compounds applied to a plurality of viable samples of a cell structure in the in vitro microscopy assay, the method comprising: receiving (212) a set of images associated with the plurality of samples; inputting (214) each image in the image set into a first ML model configured to predict a phenotypic characteristic of a cellular structure within a sample associated with said each image; inputting (216) each of the predicted phenotypic features associated with each sample into a second ML model configured to predict a lower dimensional phenotypic feature embedding for said each sample; comparing (218) the distance between the lower dimensional phenotypic feature embedding of each sample and the lower dimensional phenotypic feature embedding of the sample to which a compound with known non-toxicity or efficacy was applied; and Based on the comparison, an indication of the lack of toxicity or efficacy of each sample and the compound applied to that sample is output (220) for each sample.

13. An apparatus comprising a processor (602), a memory unit (606) and a communication interface (604), in, The processor (602) is connected to the memory unit (606) and the communication interface (604), wherein the processor (602) and the memory (606) are configured to implement the computer-implemented method as claimed in any one of the preceding claims.

14. A non-transitory tangible computer readable medium comprising data or instruction codes which, when executed on a processor (602), causes the processor (602) to implement the computer-implemented method of any one of claims 1 to 12.

15. A system, include: a sampling module (612) configured to identify a first set of samples available for analysis from a plurality of samples in an assay plate; An imager module (614) configured to generate a 2-dimensional (2D) image set for each sample in the first sample set, wherein the 2D image set of each sample includes a plurality of 2D image slices taken along a z-axis of each sample; a sample feasibility module (616) configured to identify a feasible sample set from the set of 2D image slices; as well as An output module (618) is configured to output data representing the feasible sample set for analysis.