Deep Learning Framework for Genetic Regulatory Activity Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods lack an integrated, global view of sequence regulatory activities, limiting the interpretability and effectiveness of sequence-based analysis of genomic variants and human genetics data.

Innovation Solution

A deep learning sequence model-based framework predicts comprehensive chromatin profiles for any sequence or variant, mapping sequences to regulatory activities quantitatively using a novel vocabulary of sequence classes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If current methods are used for sequence-based analysis, then analysis can be performed, but interpretability and effectiveness are limited due to lack of integrated global view

Engineering Contradiction:
ImproveinterpretabilityVSAvoidmethod complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the complex regulatory landscape into discrete sequence classes (e.g., promoters, enhancers, silencers, insulators) with specific functional characteristics. Each sequence class is defined by unique combinations of transcription factor binding sites, chromatin accessibility patterns, and epigenetic marks, enabling precise interpretation of regulatory activities without requiring analysis of the entire genome simultaneously.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces sequence classes as intermediary categories that bridge raw sequence data and regulatory function interpretation. These sequence classes serve as mediators that translate complex genomic variants into understandable regulatory outcomes by mapping variants to specific sequence class disruptions or modifications, thereby improving interpretability while managing complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If comprehensive chromatin profiles are predicted for any sequence, then global quantitative map of regulatory activities is provided, but computational model complexity increases

Engineering Contradiction:
Improveregulatory activity informationVSAvoidcomputational model complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent develops a universal deep learning model ( Sei) that can predict multiple types of chromatin profiles simultaneously from a single sequence input. The model integrates predictions for histone modifications, chromatin accessibility, transcription factor binding, and other regulatory features into a unified framework, reducing the need for multiple separate models while capturing comprehensive regulatory information.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent employs a nested architecture where the deep learning model incorporates multiple levels of feature extraction and integration. The model nests sequence features, motif patterns, chromatin states, and regulatory element predictions within a hierarchical structure, allowing comprehensive regulatory profile prediction while managing computational complexity through organized feature integration.

Inventive Principle:
Principle #7Nested doll (Nesting)

3Measurement precision

If sequences are mapped to regulatory activities quantitatively using sequence classes, then variant effects on transcriptional regulation can be predicted, but analysis complexity increases

Engineering Contradiction:
Improvevariant effect prediction accuracyVSAvoidanalysis simplicity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent applies local quality by assigning specific functional properties to different sequence classes based on their local genomic context. Each sequence class is characterized by local features such as specific transcription factor combinations, chromatin accessibility patterns, and epigenetic marks that define its regulatory function. This allows precise prediction of variant effects by evaluating changes within specific local sequence class contexts rather than requiring global genome-wide analysis.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20250182843A1Systems and Methods for Analyzing Genetic Data for Assessment of Gene Regulatory Activity
Publication Date: 2025.06.05 THE SIMONS FOUND INC
  • US20250182843A1 patent drawing
  • US20250182843A1 patent drawing
  • US20250182843A1 patent drawing

AI summary

Processes that determine transcriptional regulation from genetic sequence data are described. Generally, computational models are trained to predict transcriptional regulatory effects, which can be used in several downstream applications. Various methods further develop research tools, develop and perform diagnostics, and treat individuals based on identified variants.