PerturbNet AI for Predicting Unseen Cell State Transitions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Deep generative models struggle to predict gene expression of unseen cell states and are limited in their ability to generate realistic high-dimensional single-cell data, constraining their applicability due to their reliance on observed conditions during training.

Innovation Solution

A computer-implemented method and system, such as PerturbNet, that uses a trained machine learning algorithm to identify perturbations causing a starting cell state to transition to a target cell state by encoding chemical and genetic perturbations into latent representations, enabling predictions for both observed and unseen drug treatments and genetic perturbations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If deep generative models are trained on observed cell states, then they can generate realistic data similar to training data, but they fail to predict gene expression of unseen cell states

Engineering Contradiction:
Improveprediction accuracy for unseen cell statesVSAvoidability to generalize to unseen conditions
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent transforms discrete cell state labels into continuous latent representations using autoencoders, enabling the model to generalize to unseen states by interpolating in the continuous latent space rather than relying on discrete observed categories

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces a latent representation dimension that bridges observed and unseen cell states, allowing predictions in previously unobserved regions of the cell state space through continuous transformation rather than discrete classification

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Manufacturing precision

If deep generative models focus on generating realistic data similar to training data, then they achieve good reconstruction quality, but they have limited ability to predict gene expression of unseen cell states

Engineering Contradiction:
Improvedata generation qualityVSAvoidprediction capability for unseen states
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The patent separates the model into distinct components: an autoencoder for learning latent representations and a predictor for generating gene expression predictions, allowing each component to optimize for its specific function while working together

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The latent representation serves as an intermediary between observed cell states and unseen predictions, encoding essential features in a compressed form that enables generalization to new conditions while preserving generation quality

Inventive Principle:
Principle #24Intermediary (Mediator)

3Quantity of substance

If single-cell data has a relatively small set of observed conditions, then training data is limited, but this constrains the applicability of deep generative models

Engineering Contradiction:
Improvenumber of observed conditionsVSAvoidmodel applicability range
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent transforms the limited discrete observed conditions into a continuous latent space, enabling the model to generate predictions for any point in the continuous space, effectively expanding the applicability beyond the limited training conditions

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240047081A1Designing Chemical or Genetic Perturbations using Artificial Intelligence
Publication Date: 2024.02.08 THE RGT UNIV OF MICHIGAN
  • US20240047081A1 patent drawing
  • US20240047081A1 patent drawing
  • US20240047081A1 patent drawing

AI summary

The following relates generally to identifying perturbations (e.g., chemical perturbations, or genetic perturbations), drug treatments, and/or protein sequences. Some embodiments include a machine learning algorithm comprising a first network that converts perturbations into real-valued vector representations of the perturbations; a second network that converts cell states into real-valued vector representations of the cell states; and a third network that maps relationships between: (i) the real-valued vector representations of the perturbations, and (ii) the real-valued vector representations of the cell states. Some embodiments use the machine learning algorithm to identify a perturbation that will cause a starting cell state to transition to a target cell state by inputting the starting cell state and the target cell state into the machine learning algorithm.