Clinical-Phenomics Causal Discovery for Gene-Outcome Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems face inaccuracies, inefficiencies, and operational inflexibilities in mapping relationships between clinical data and specific diseases due to high dimensionality and noise in clinically observed data, requiring excessive computational resources and time.

Innovation Solution

A causal discovery system utilizing phenomic image embeddings, clustering models, machine learning classification models, and causal discovery models to generate accurate and efficient causal predictions between genes and clinical outcomes by filtering data biologically and flexibly combining phenomic and clinical observation data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional systems utilize computational models to parse through clinical data and map clinical features to specific diseases, then relationships between health abnormalities and clinically observed factors can be determined, but the systems experience inaccuracies and require excessive computational resources and time due to high dimensionality and noise in clinically observed data

Engineering Contradiction:
Improveaccuracy of mapping relationshipsVSAvoidcomputational resources required
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the complex clinical data analysis into distinct modules: a phenotyping module that processes phenomic data to identify phenotypic patterns, a gene prioritization module that ranks genes based on phenotypic associations, and a causal inference module that determines causal relationships. This segmentation reduces computational complexity by breaking down the high-dimensional data processing into manageable stages, each handling specific aspects of the analysis independently.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces phenomic data as an intermediary layer between clinical observations and gene-disease mappings. Phenomic data serves as a mediator that captures intermediate phenotypic states, reducing the direct complexity of mapping clinical features to diseases. This intermediary approach filters and structures data before final causal inference, decreasing the computational burden while improving accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If conventional systems parse through large volumes of clinical data to map clinical features to specific diseases, then relationships between health abnormalities and clinically observed factors can be identified, but the process experiences inefficiencies and time consumption

Engineering Contradiction:
Improveaccuracy of disease mappingVSAvoidefficiency of data processing
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent performs preliminary phenotyping analysis before conducting full disease mapping. The phenotyping module pre-processes phenomic data to identify and structure phenotypic patterns in advance, creating organized intermediate representations that accelerate subsequent gene prioritization and causal inference steps. This preliminary structuring of data significantly reduces processing time for the remaining analysis stages.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements dynamic gene prioritization that adapts to the specific phenotypic patterns observed in the data. Rather than using static prioritization methods, the system dynamically adjusts gene ranking based on phenotypic associations, allowing the analysis to focus computational resources on the most relevant gene-disease relationships identified during phenotyping, thereby improving processing efficiency.

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If conventional systems map clinical features to specific diseases using computational models, then relationships between health abnormalities and clinically observed factors can be determined, but the systems experience operational inflexibilities

Engineering Contradiction:
Improveaccuracy of relationship mappingVSAvoidoperational flexibility
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal phenotyping framework that can handle multiple data types and disease categories through a common phenotypic representation system. The phenotyping module is designed to process various phenomic data formats and generate standardized phenotypic patterns that can be applied across different disease contexts, enabling the system to adapt to diverse analytical needs while maintaining consistent accuracy standards.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent enables flexible parameter adjustment in the gene prioritization and causal inference modules, allowing users to modify prioritization thresholds, phenotypic association criteria, and causal inference parameters based on specific research questions or data characteristics. This parameter flexibility allows the system to adapt its operational characteristics while maintaining accurate relationship mapping across different operational contexts.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250378905A1Utilizing a clinical-phenomics causal discovery framework to generate causal discovery predictions
Publication Date: 2025.12.11 RECURSION PHARMACEUTICALS INC
  • US20250378905A1 patent drawing
  • US20250378905A1 patent drawing
  • US20250378905A1 patent drawing

AI summary

The present disclosure relates to systems, non-transitory computer-readable media, and methods that analyze gene perturbation machine learning embeddings and clinical observation data sets utilizing machine learning, explainability models, and causal discovery models to generate causal predictions between one or more genes and clinical outcomes. Indeed, in one or more implementations, the disclosed systems identify gene perturbation embeddings generated from cells exposed to perturbations. For instance, the disclosed systems select a cluster of genes from a plurality of genes by applying a clustering model to the gene perturbation embeddings. In some instances, the disclosed systems select gene targets from the cluster of genes by using a machine learning classification model trained on a plurality of features of the clinical observation data set. Moreover, in some instances, the disclosed systems generate the causal prediction from the gene targets and the clinical observation data set utilizing a causal discovery model.