Cellular Toxicity Prediction Using Phenotype Embedding Distance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional semi-automated toxicity prediction assays using High Throughput Screening (HTS) are unreliable in identifying the toxicity of compounds, leading to increased risks of drug-induced liver injury during in-vivo trials, with up to 20-40% of drug-induced liver injury cases presenting cholestatic and/or mixed hepatocellular/cholestatic patterns.

Innovation Solution

A computer-implemented method using a combination of machine learning models, including a neural network for phenotype feature extraction and a Uniform Manifold Approximation and Projection (UMAP) algorithm for dimensional reduction, to predict toxicity by comparing the distance between cellular structure embeddings and known toxic compounds, enhancing the reliability of in-vitro toxicity prediction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional semi-automated HTS assays are used for toxicity prediction, then large numbers of compounds can be screened quickly, but the reliability of toxicity identification is insufficient

Engineering Contradiction:
Improvescreening speedVSAvoidtoxicity identification accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent replaces conventional image microscopy and manual DR graph analysis with a deep learning-based computational system. The system uses neural networks to process microscopy images and automatically predict toxicity, substituting mechanical/optical measurement systems with intelligent algorithms that analyze cellular phenotypes and generate toxicity predictions without manual intervention.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent transforms the toxicity assessment from qualitative DR graph analysis to quantitative phenotype embedding distance measurements. By converting cellular responses into numerical embeddings and calculating distances from control samples, the system changes the measurement parameters from visual dose-response curves to computable metric spaces, enabling more reliable and automated toxicity classification.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If deep learning models are used for phenotype feature extraction and dimensional reduction, then toxicity prediction reliability is improved, but system complexity increases

Engineering Contradiction:
Improvetoxicity prediction accuracyVSAvoidmodel architecture complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent divides the deep learning system into two distinct modules: a first model for phenotype feature extraction from microscopy images, and a second model for dimensional reduction to create phenotype embeddings. This segmentation allows each model to specialize in one task, improving overall reliability while making the complex system more manageable through functional decomposition.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces phenotype embeddings as an intermediary representation between raw microscopy images and final toxicity predictions. The embedding layer transforms high-dimensional image features into a compressed, informative latent space that captures essential cellular phenotypes, serving as a bridge that simplifies the transition from complex image data to interpretable toxicity metrics.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250356959A1Toxicity Prediction Of Compounds In Cellular Structures
Publication Date: 2025.11.20 SANOFI SA(FR)
  • US20250356959A1 patent drawing
  • US20250356959A1 patent drawing
  • US20250356959A1 patent drawing

AI summary

Methods, apparatus and systems are disclosed for predicting toxicity of one or more compounds applied to a plurality of samples of a cellular structure in an in-vitro microscopy assay. A set of images associated with the plurality of samples are received. Each image of the set of images are input to a first ML model configured for predicting phenotype features of the cellular structure within the sample associated with said each image. Each of the predicted phenotype features associated with each sample are input to a second ML model configured for predicting a lower dimensional phenotype feature embedding of said each sample. The distance between the lower dimensional phenotype feature embedding of said each sample is compared with that of a sample applied with a compound having a known toxicity. For each sample, an indication of the toxicity of said each sample and applied compound thereto is output based on said comparison.