Contextual DNA Sequence Modeling for Tissue-Specific Gene Expression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current gene editing technologies, such as CRISPR/Cas9, primarily disrupt coding sequences for genetic enhancement, failing to effectively fine-tune noncoding regulatory sequences for gene expression patterns, which are crucial for achieving genetic gains in crops.
Innovation Solution
A transformer-based machine learning model is trained to predict mRNA and/or protein expression across different types of tissues using a transformer-based machine learning model to predict mRNA and/or protein expression across different types of tissues of an organism, identifying regulatory regions via saliency scores.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If CRISPR/Cas9 is used to disrupt coding sequences for genetic enhancement, then large changes in gene expression can be achieved, but the ability to fine-tune noncoding regulatory sequences for precise expression patterns is lost
Solution Approach 1:
The patent introduces an AI model as an intermediary between researchers and noncoding regulatory sequences. The model predicts tissue-specific expression patterns and identifies functional regulatory elements, enabling precise targeting of editing efforts to the most impactful noncoding regions without requiring exhaustive experimental screening of all potential regulatory elements.
Solution Approach 2:
The patent performs preliminary computational analysis of noncoding regulatory sequences using the AI model before conducting physical gene editing. This preliminary action identifies which noncoding regions are most likely to influence tissue-specific expression, allowing researchers to prioritize editing efforts and achieve fine-tuned expression control more efficiently.
2Reliability
If conserved noncoding sequences are identified across species, then functionally constrained regulatory regions can be detected, but evolutionarily young regulatory regions and species-specific regulatory elements are missed
Solution Approach 1:
The patent employs a dynamic AI model that adapts to different species and regulatory contexts rather than relying on static conservation thresholds. The model learns tissue-specific expression patterns and regulatory relationships specific to each species, enabling reliable identification of both conserved and species-specific regulatory elements through flexible, context-dependent predictions.
Solution Approach 2:
The patent changes the fundamental parameter for identifying regulatory sequences from sequence conservation to predicted tissue-specific expression impact. This parameter change allows the system to detect regulatory elements based on their functional consequence on gene expression patterns, capturing both ancient conserved elements and newer species-specific elements that may not show sequence conservation but have clear regulatory effects.
3Quantity of substance
If open chromatin regions are identified using ATAC-seq or MNase-seq, then a superset of regulatory sequences is obtained, but the precise functional regulatory elements within these regions cannot be distinguished
Solution Approach 1:
The patent replaces the mechanical/experimental approach of chromatin accessibility assays (ATAC-seq, MNase-seq) with an AI-based computational system. Instead of physically probing chromatin accessibility to identify potential regulatory regions, the model computationally predicts which noncoding sequences have the greatest impact on tissue-specific expression patterns, directly identifying functional elements rather than just accessible regions.
Solution Approach 2:
The patent extracts the specific functional signal from within the broad category of open chromatin regions. By training the AI model to predict tissue-specific expression patterns, the system extracts and identifies only those noncoding sequences that actually function as regulatory elements, separating the functionally active elements from the larger set of merely accessible chromatin regions.
4Device complexity
If proximal promoter sequences are used for prediction, then the model complexity is reduced, but the ability to predict tissue-specific expression patterns accurately is compromised
Solution Approach 1:
The patent extends the analysis from one dimension (proximal promoter sequences only) to multiple dimensions by incorporating distal noncoding sequences, enhancers, and broader genomic contexts into the prediction model. This dimensional expansion allows the model to capture long-range regulatory interactions and tissue-specific expression patterns that cannot be predicted from proximal promoters alone, improving accuracy without being constrained by model simplicity.
Data Source
AI summary
Systems and methods are presented for constructing, training, and utilizing a large contextual gene sequence model that can be provided with variable-length DNA sequence data from larger genomic intervals surrounding annotated genes to predict relative expression across a set of transcriptionally diverse tissues. Additionally, per-nucleotide saliency scores can be extracted from the large contextual gene sequence model, indicating which regions of the DNA sequence surrounding a target gene are associated with regulation of expression of the target gene in various tissue types of an organism. The model described herein surprisingly tolerated averaging across a given embedding of a DNA sequence, and the use of averaging across each embedding allowed the development of a transformer-based model capable of producing a constant set of outputs, indicating the relative expression across a fixed set of tissues, from variable lengths of DNA sequence.


