AI Biomarker Screening From Whole-Slide Tile Embeddings

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional molecular biomarker assays for cancer are expensive, tissue-destructive, and time-consuming, and lack systems for pan-cancer and pan-tissue analysis, failing to efficiently utilize large datasets across various modalities.

Innovation Solution

A high-throughput AI-based system using a foundation model trained on millions of whole slide images to predict a wide range of molecular biomarkers across cancer types, leveraging a unified model to aggregate tile-level embeddings into slide-level predictions, incorporating an attention mechanism for efficient biomarker prediction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional molecular assays are used to detect biomarkers, then measurement precision is maintained, but productivity is low and loss of time is high

Engineering Contradiction:
Improvebiomarker detection accuracyVSAvoidbiomarker screening throughput
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The digital medical image is divided into multiple tiles, and the foundation model analyzes each tile independently to generate embedding vectors. This segmentation enables parallel processing of multiple regions simultaneously, significantly increasing throughput while maintaining detection accuracy through comprehensive coverage of the entire slide.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The foundation model creates embedding vector representations (digital copies) of tissue tiles that capture essential features without requiring physical tissue consumption. These embedding vectors serve as efficient proxies for the actual tissue data, enabling rapid analysis and multiple predictions from the same source material.

Inventive Principle:
Principle #26Copying

2Measurement precision

If conventional molecular assays are used, then measurement precision is maintained, but loss of time is high

Engineering Contradiction:
Improvebiomarker detection accuracyVSAvoidturnaround time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The foundation model is pre-trained on millions of whole slide images before deployment. This preliminary training equips the model with comprehensive knowledge of various tissue patterns and biomarker characteristics, enabling it to perform accurate predictions rapidly without requiring time-consuming retraining or analysis for each new case.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Embedding vectors serve as compressed digital representations that capture essential tissue information in a compact format. These vectors enable rapid aggregation and analysis by the aggregator model, significantly reducing the time required to process and interpret tissue data while preserving measurement precision.

Inventive Principle:
Principle #26Copying

3Measurement precision

If individual models are trained for each biomarker or cancer type, then measurement precision is improved, but device complexity and loss of substance increase

Engineering Contradiction:
Improvebiomarker prediction accuracyVSAvoidmodel training complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The foundation model is designed as a universal system trained on millions of diverse whole slide images across multiple cancer types. This single multi-functional model can predict multiple different biomarkers across various cancer types, eliminating the need for separate specialized models while maintaining measurement precision through its comprehensive training data and transfer learning capabilities.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system merges the foundation model's tile-level prediction capabilities with the aggregator model's slide-level synthesis functions into a unified pan-cancer prediction system. This combination integrates multiple functions (tile analysis, embedding generation, attention-based aggregation, and biomarker prediction) into a single cohesive workflow, reducing overall system complexity while improving comprehensive accuracy.

Inventive Principle:
Principle #5Merging (Combining)

4Measurement precision

If conventional assays are used, then measurement precision is maintained, but productivity is low

Engineering Contradiction:
Improvegenomic abnormality detection accuracyVSAvoidpan-cancer screening throughput
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system segments the analysis into tile-level operations using the foundation model, which processes individual image tiles in parallel. This segmentation enables high-throughput processing of entire tissue slides by distributing computational work across multiple independent tile analyses, significantly increasing pan-cancer screening throughput while maintaining detection accuracy through aggregate analysis of all tiles.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The foundation model generates embedding vector copies that efficiently represent tissue tile information. These compact vector representations enable rapid aggregation and comparison across multiple tiles and cancer types, facilitating high-throughput pan-cancer genomic abnormality detection while preserving measurement precision through the rich information contained in the embedding vectors.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20260066122A1Systems and methods for high-throughput pan-cancer genetic and phenotypic biomarker screening
Publication Date: 2026.03.05 PAIGE AI INC
  • US20260066122A1 patent drawing
  • US20260066122A1 patent drawing
  • US20260066122A1 patent drawing

AI summary

Disclosed are systems and methods for processing at least one digital medical image to predict a first biomarker, including receiving the at least one digital medical image of one or more tissues of a patient, the at least one digital medical image including a plurality of tiles, analyzing, via a foundation model, the plurality of tiles to determine an embedding vector for each of the plurality of tiles, the foundation model having been trained to predict embedding vectors at a tile-level based on a plurality of digital medical images, and analyzing, via an aggregator model, the embedding vector for each of the plurality of tiles to predict the first biomarker of the digital medical image, wherein the aggregator model includes an attention mechanism configured to aggregate the embedding vector for each of the plurality of tiles into at least one slide-level prediction.