Machine-Learning Perturbation Embeddings for Cross-Experiment Heatmaps

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems for analyzing biological relationships in digital data are inaccurate, inefficient, and inflexible, failing to accurately relate experimental data from disparate perturbation experiments and requiring excessive user interactions and queries to identify subtle relationships.

Innovation Solution

A perturbation mapping system that utilizes machine learning models to embed, filter, align, and aggregate phenomic digital images of cellular perturbations, generating a genome-wide perturbation database for real-time analysis and interactive heatmaps to display perturbation relationships.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional systems use brute-force approach to search and analyze digital biological data, then they can generate a large volume of user interfaces for data analysis, but they suffer from inaccuracy, inefficiency, and operational inflexibility in utilizing large digital data volumes

Engineering Contradiction:
Improvevolume of user interfaces generatedVSAvoidaccuracy of biological relationship analysis
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent replaces the mechanical brute-force search approach with a machine learning-based embedding system. The system uses neural network models to embed phenomic digital images into low-dimensional spaces, automatically identifying biological relationships without exhaustive searching. This substitution of mechanical computation with intelligent algorithms resolves the contradiction by maintaining high productivity while significantly improving analysis accuracy through learned patterns rather than random sampling.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent transforms the data representation parameters by converting high-dimensional phenomic images into low-dimensional embeddings. This parameter transformation enables efficient comparison and analysis of biological relationships. By changing the dimensional parameters and using similarity metrics in the embedded space, the system achieves both high productivity in generating analysis results and high reliability in identifying accurate biological relationships.

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If conventional systems perform exhaustive search analysis on digital repositories, then they can present query results for display, but they require excessive user interactions and queries to identify subtle relationships

Engineering Contradiction:
Improveuser interactions requiredVSAvoidtime to identify relationships
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-computing embeddings for all perturbations and storing them in a database. When a user queries the system, the embeddings are already prepared and can be rapidly compared using similarity metrics. This preliminary processing eliminates the need for exhaustive search at query time, dramatically reducing both the number of user interactions needed and the time required to identify subtle relationships while maintaining ease of operation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements a dynamic query response system where the interface can generate perturbation heatmaps and similarity analyses in real-time based on user needs. The system dynamically adapts to different query types and can adjust the level of detail in results, reducing the number of iterative user interactions required while maintaining operational ease and minimizing time loss.

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If conventional systems use pre-generated data tables and curated lists for displaying biological data, then they can present search results, but they lack flexibility in analyzing large digital data volumes across computer networks

Engineering Contradiction:
Improveflexibility in data analysisVSAvoidsystem architecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal embedding framework that can handle multiple types of biological data (phenomic images, perturbation data, genomic data) through a single machine learning model architecture. This universal system can be deployed across computer networks and adapts to different analysis needs without requiring separate specialized systems, thereby increasing flexibility while managing complexity through a unified approach rather than multiple specialized components.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces an intermediary embedding layer that mediates between raw phenomic data and various analysis queries. This intermediary representation in low-dimensional space serves as a universal interface that can efficiently respond to different types of biological relationship queries. The embedding acts as a mediator that simplifies complex data relationships while maintaining the flexibility to answer diverse analytical questions across networked systems.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12374429B1Utilizing machine learning models to synthesize perturbation data to generate perturbation heatmap graphical user interfaces
Publication Date: 2025.07.29 RECURSION PHARMACEUTICALS INC
  • US12374429B1 patent drawing
  • US12374429B1 patent drawing
  • US12374429B1 patent drawing

AI summary

The present disclosure relates to systems, non-transitory computer-readable media, and methods for embedding perturbation data via a machine learning model and filtering, aligning, and aggregating the embeddings to generate a genome-wide perturbation database for real-time generation of perturbation heatmaps. In particular, in one or more embodiments, the disclosed systems can receive a plurality of perturbation images portraying cells from a plurality of wells corresponding to a plurality of cell perturbations. Further, the systems can generate, utilizing a machine learning model, a plurality of well-level image embeddings from the plurality of perturbation images. Moreover, the systems can align, utilizing an alignment model, the plurality of well-level image embeddings to generate aligned well-level image embeddings. Additionally, the systems can aggregate, according to perturbations of one or more perturbation experiments, the well-level image embeddings to generate perturbation-level image embeddings. Furthermore, the systems can generate perturbation comparisons utilizing the perturbation-level image embeddings.