Masked Autoencoder Embeddings for Microscopy Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems for analyzing digital signals from microscopy images face challenges such as computational inefficiencies, extraction inaccuracies, and inflexibilities in training and utilizing machine learning models.

Innovation Solution

The use of generative machine learning models, specifically masked autoencoder generative models, to generate embeddings from phenomic images, enabling efficient training and analysis without the need for segmentation or classification labels.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional deep vision models are used to extract features from microscopy images, then biological relationships can be inferred from cellular phenotypes, but computational inefficiencies and repetitive training iterations occur

Engineering Contradiction:
Improveextraction accuracyVSAvoidtraining efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs preliminary action by pre-training the masked autoencoder model on large batches of microscopy images to learn robust feature representations before actual analysis. This pre-training phase captures general biological patterns that can be reused across different experiments, eliminating the need for repetitive full-model training on each new dataset while maintaining high extraction accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The invention extracts and separates the feature extraction function from the complete classification pipeline by using a masked autoencoder that learns representations independently. This extracted embedding layer can be reused across multiple tasks and datasets, improving computational efficiency while preserving the accuracy needed for inferring biological relationships.

Inventive Principle:
Principle #2Taking out (Extraction)

2Measurement precision

If large-scale training batches are used to improve model accuracy, then extraction precision improves, but computational resources and GPU time increase

Engineering Contradiction:
Improveembedding accuracyVSAvoidGPU time
Core Design Contradiction:
Measurement precisionVSUse of energy by stationary object

Solution Approach 1:

The system performs a single large-scale pre-training pass on comprehensive datasets to establish accurate feature embeddings, then reuses these pre-trained weights for subsequent analyses. This preliminary action captures the benefits of large-batch training once, avoiding repeated consumption of computational resources while maintaining high embedding accuracy across multiple experiments.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The invention creates reusable embedding representations from large-scale training that can be copied and applied to multiple different datasets and analysis tasks. Instead of retraining on each new dataset, the system copies the learned feature extraction capabilities, dramatically reducing GPU time and energy consumption while preserving the accuracy benefits of large-scale training.

Inventive Principle:
Principle #26Copying

3Measurement precision

If conventional models require segmentation or classification labels for training, then supervised learning accuracy can be achieved, but training flexibility and ease of operation decrease

Engineering Contradiction:
Improvesupervised learning accuracyVSAvoidtraining flexibility
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The masked autoencoder model performs self-service by learning feature representations in an unsupervised manner from raw microscopy images without requiring external segmentation or classification labels. The model serves its own training needs by reconstructing input images and learning from reconstruction errors, eliminating the operational burden of label preparation while still achieving accurate biological feature extraction.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The invention creates a universal feature extraction model that can operate across multiple different microscopy datasets and biological questions without requiring dataset-specific labeled training data. This multi-functional approach maintains the accuracy benefits of supervised learning by learning general biological patterns that transfer across tasks, while dramatically improving training flexibility and ease of operation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250201351A1Utilizing masked autoencoder generative models to extract microscopy representation autoencoder embeddings
Publication Date: 2025.06.19 RECURSION PHARMACEUTICALS INC
  • US20250201351A1 patent drawing
  • US20250201351A1 patent drawing
  • US20250201351A1 patent drawing

AI summary

The present disclosure relates to systems, non-transitory computer-readable media, and methods for training and utilizing generative machine learning models to generate embeddings from phenomic images (or other microscopy representations). For example, the disclosed systems can train a generative machine learning model (e.g., a masked autoencoder generative model) to generate predicted (or reconstructed) phenomic images from masked version of ground truth training phenomic images. In some cases, the disclosed systems utilize a momentum-tracking optimizer while reducing a loss of the generative machine learning model to enable efficient training on large scale training image batches. Furthermore, the disclosed systems can utilize Fourier transformation losses with multi-stage weighting to improve the accuracy of the generative machine learning model on the phenomic images during training. Indeed, the disclosed systems can utilize the trained generative machine learning model to generate phenomic embeddings from input phenomic images (for various phenomic comparisons).