AI Epigenetic Chromatin Modeling for Base-Resolution Variant Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing genomics data analysis methods struggle to accurately identify genetic variants causing extreme levels of gene expression due to confounding factors and incomplete data representation, leading to reduced classification accuracy and difficulty in diagnosing pathogenicity of genetic diseases.

Innovation Solution

An artificial intelligence-based approach that utilizes a chromatin model to process genomic data at base resolution, incorporating epigenetic signals to identify rare variants causing extreme gene expression levels by controlling for confounders using a causality model and generating causality scores.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing genomics data analysis methods are used, then analysis can be performed with available tools, but classification accuracy is reduced due to confounding factors and incomplete data representation

Engineering Contradiction:
Improveclassification accuracyVSAvoididentifying genetic variants
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent segments the analysis into multiple independent components: chromatin accessibility data, histone modification data, DNA sequence features, and gene expression data are processed separately through distinct neural network branches before being integrated. This segmentation allows each feature type to be optimized independently while reducing confounding effects between different data modalities.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a composite analytical model that integrates multiple types of genomic data (chromatin accessibility, histone modifications, DNA sequences, and gene expression) into a unified deep learning framework. This composite approach combines the strengths of different data types to achieve higher classification accuracy than any single data type alone.

Inventive Principle:
Principle #40Composite materials

2Measurement precision

If deep learning models integrate multiple data types, then classification accuracy improves, but model complexity increases

Engineering Contradiction:
Improveclassification accuracyVSAvoidmodel complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The deep learning model is divided into separate processing branches for different data types (chromatin accessibility, histone modifications, DNA sequences, gene expression), with each branch specialized for its input type. This modular segmentation reduces training complexity while maintaining the benefits of multi-modal integration.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediate representation layers that transform raw genomic data into standardized feature vectors before final integration. These intermediary layers act as adapters between different data modalities, simplifying the integration process and reducing overall model complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If rare variants are identified with higher sensitivity, then diagnostic capabilities improve, but false positives increase due to confounding factors

Engineering Contradiction:
Improvesensitivity of identifying genetic variantsVSAvoiddiagnostic accuracy
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent combines multiple independent lines of evidence (chromatin accessibility, histone modifications, DNA sequence conservation, and gene expression patterns) into a composite scoring system. Rare variants are only flagged as high-confidence when they show consistent signals across multiple data types, thereby increasing sensitivity while reducing false positives through cross-validation.

Inventive Principle:
Principle #40Composite materials

Solution Approach 2:

The model incorporates feedback mechanisms where predictions from one data type inform the weighting and interpretation of other data types. Confounding factors detected in one modality can suppress signals in other modalities, creating a self-correcting system that reduces false positives while maintaining sensitivity to true rare variants.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250218534A1Artificial intelligence-based epigenetics at base resolution
Publication Date: 2025.07.03 ILLUMINA INC
  • US20250218534A1 patent drawing
  • US20250218534A1 patent drawing
  • US20250218534A1 patent drawing

AI summary

The technology disclosed relates to reliably identifying variants that cause extreme levels of gene expression. Extreme levels of gene expression include under expression and over expression. Then, these variants are used to train artificial intelligence based models for a variety of prediction tasks. One example of the prediction tasks is to produce per-base resolution for chromatin sequences. Another example of the chromatin task is to produce gene expression changes caused by the reliably identified variants.