Contextual DNA Sequence Modeling for Tissue-Specific Gene Expression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current gene editing technologies, such as CRISPR/Cas9, primarily disrupt coding sequences for genetic enhancement, failing to effectively fine-tune noncoding regulatory sequences for gene expression patterns, which are crucial for achieving genetic gains in crops.

Innovation Solution

A transformer-based machine learning model is trained to predict mRNA and/or protein expression across different types of tissues using a transformer-based machine learning model to predict mRNA and/or protein expression across different types of tissues of an organism, identifying regulatory regions via saliency scores.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If CRISPR/Cas9 is used to disrupt coding sequences for genetic enhancement, then large changes in gene expression can be achieved, but the ability to fine-tune noncoding regulatory sequences for precise expression patterns is lost

Engineering Contradiction:
Improveprecision of gene expression controlVSAvoidversatility of gene editing approaches
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The patent introduces an AI model as an intermediary between researchers and noncoding regulatory sequences. The model predicts tissue-specific expression patterns and identifies functional regulatory elements, enabling precise targeting of editing efforts to the most impactful noncoding regions without requiring exhaustive experimental screening of all potential regulatory elements.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent performs preliminary computational analysis of noncoding regulatory sequences using the AI model before conducting physical gene editing. This preliminary action identifies which noncoding regions are most likely to influence tissue-specific expression, allowing researchers to prioritize editing efforts and achieve fine-tuned expression control more efficiently.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If conserved noncoding sequences are identified across species, then functionally constrained regulatory regions can be detected, but evolutionarily young regulatory regions and species-specific regulatory elements are missed

Engineering Contradiction:
Improvereliability of regulatory region identificationVSAvoidcoverage of diverse regulatory sequences
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent employs a dynamic AI model that adapts to different species and regulatory contexts rather than relying on static conservation thresholds. The model learns tissue-specific expression patterns and regulatory relationships specific to each species, enabling reliable identification of both conserved and species-specific regulatory elements through flexible, context-dependent predictions.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the fundamental parameter for identifying regulatory sequences from sequence conservation to predicted tissue-specific expression impact. This parameter change allows the system to detect regulatory elements based on their functional consequence on gene expression patterns, capturing both ancient conserved elements and newer species-specific elements that may not show sequence conservation but have clear regulatory effects.

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If open chromatin regions are identified using ATAC-seq or MNase-seq, then a superset of regulatory sequences is obtained, but the precise functional regulatory elements within these regions cannot be distinguished

Engineering Contradiction:
Improvenumber of identified regulatory regionsVSAvoidprecision of functional regulatory element identification
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent replaces the mechanical/experimental approach of chromatin accessibility assays (ATAC-seq, MNase-seq) with an AI-based computational system. Instead of physically probing chromatin accessibility to identify potential regulatory regions, the model computationally predicts which noncoding sequences have the greatest impact on tissue-specific expression patterns, directly identifying functional elements rather than just accessible regions.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent extracts the specific functional signal from within the broad category of open chromatin regions. By training the AI model to predict tissue-specific expression patterns, the system extracts and identifies only those noncoding sequences that actually function as regulatory elements, separating the functionally active elements from the larger set of merely accessible chromatin regions.

Inventive Principle:
Principle #2Taking out (Extraction)

4Device complexity

If proximal promoter sequences are used for prediction, then the model complexity is reduced, but the ability to predict tissue-specific expression patterns accurately is compromised

Engineering Contradiction:
Improvecomplexity of prediction modelVSAvoidaccuracy of tissue-specific expression prediction
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent extends the analysis from one dimension (proximal promoter sequences only) to multiple dimensions by incorporating distal noncoding sequences, enhancers, and broader genomic contexts into the prediction model. This dimensional expansion allows the model to capture long-range regulatory interactions and tissue-specific expression patterns that cannot be predicted from proximal promoters alone, improving accuracy without being constrained by model simplicity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20250372209A1Systems and methods for identifying DNA sequences regulating pattern of expression for genes of interest
Publication Date: 2025.12.04 NUTECH VENTURES LTD
  • US20250372209A1 patent drawing
  • US20250372209A1 patent drawing
  • US20250372209A1 patent drawing

AI summary

Systems and methods are presented for constructing, training, and utilizing a large contextual gene sequence model that can be provided with variable-length DNA sequence data from larger genomic intervals surrounding annotated genes to predict relative expression across a set of transcriptionally diverse tissues. Additionally, per-nucleotide saliency scores can be extracted from the large contextual gene sequence model, indicating which regions of the DNA sequence surrounding a target gene are associated with regulation of expression of the target gene in various tissue types of an organism. The model described herein surprisingly tolerated averaging across a given embedding of a DNA sequence, and the use of averaging across each embedding allowed the development of a transformer-based model capable of producing a constant set of outputs, indicating the relative expression across a fixed set of tissues, from variable lengths of DNA sequence.